오픈AI가 아직 내놓지도 않은 모델에 제동을 걸었습니다. 8월 7일 발표입니다.

위험하다고 확인해서가 아닙니다. 위험하지 않다고 말할 수 없어서입니다.

결과가 아니라 판단 중인 상태를 먼저 내놓았다는 점이 이번 발표의 성격입니다.

■ '중대'가 무슨 뜻인지부터

'중대'가 무슨 뜻일까요? 오픈AI에는 준비 프레임워크라는 자체 기준이 있습니다. 모델 능력이 어느 선을 넘으면 회사가 무엇을 할지 미리 정해둔 문서이고, 2023년 12월에 처음 공개됐습니다.

여기서 사이버보안 '중대(Critical)' 등급은 이렇게 정의돼 있습니다. 사람이 개입하지 않고도 견고한 실제 핵심 시스템 다수에서 모든 등급의 제로데이 취약점을 찾아 실제 동작하는 공격 코드를 만들어내는 수준. 또는 높은 수준의 목표만 주면 견고한 표적에 대한 새로운 공격 전략을 처음부터 끝까지 세우고 실행하는 수준.

오픈AI는 곧 내놓을 모델 아스트라를 며칠간 내부 평가했습니다. 에이전트 코딩과 사이버보안 능력이 크게 올라간 것으로 나왔습니다.

그 결과와 전문가 평가를 합쳐, 아스트라가 이 중대 등급에 해당하지 않는다고 단정할 수 없다는 결론에 이르렀습니다. 발표문은 그 시점을 “지난밤”이라고 적었습니다.

■ 확정을 기다리지 않고 걸었다

확정도 안 된 상태에서 무엇을 했을까요? 오픈AI가 밝힌 내부 조치는 이렇습니다. 격리된 테스트 환경과 네트워크·도구 접근 제한, 모델 가중치 보호와 암호화 강화, 모니터링과 탐지 추가, 샌드박스 실행.

강화된 보안 요건을 아직 못 맞춘 아스트라 관련 내부 작업은 중단했습니다. 학습과 평가를 포함한 모든 에이전트 활용에는 위험 행동과 정렬 이탈을 상시 감시하고, 모델의 사고 사슬을 보며 고위험 활동을 검토하거나 중단합니다.

검증에는 정부 기관과 일부 AI 안전 기관을 들이고, 외부 테스트 파트너에게는 고위험 평가를 안전하게 돌릴 권고 보안 통제를 제공한다고 했습니다. 안전장치 자체에 대한 견고성 시험도 이 능력 수준에 맞게 끌어올렸다고 밝혔습니다.

하나 못 박은 것도 있습니다. 아스트라는 허깅페이스 침해에 관여하지 않았다는 대목입니다.

■ 이게 처음은 아니다

오픈AI는 2025년 6월에도 같은 방식으로 움직였습니다. 당시 모델들이 생물학 고역량 임계값에 가까워지자 안전장치를 강화하고 시험을 넓히고 외부 전문가와 협력하고 보안 통제를 추가했습니다.

이번에도 같은 원칙을 적용했다는 설명입니다. 프레임워크를 만든 이유가 이런 순간에 무엇을 할지 미리 정해두기 위해서였다는 것입니다. GPT-5.6 솔을 비롯한 이전 모델들도 최전선 사이버 역량 평가를 받아왔습니다.

■ 같은 능력이 이미 쓰이고 있다

그런데 이건 아직 안 나온 모델 이야기일 뿐일까요? 비슷한 능력을 가진 모델이 이미 국가 기관에서 돌아가고 있습니다.

미국 국가안보국(NSA)은 앤트로픽의 사이버 AI 모델 '미토스'를 '실험적'으로 계속 쓰고 있습니다. 미군 네트워크의 방어력을 시험하고, 중국·러시아·북한·이란 네트워크의 잠재적 취약점을 찾는 데 씁니다. 최근 정부를 떠난 한 고위 관리는 미토스 사용을 중단하는 것이 사이버 경쟁에서 '일방적 무장해제'에 해당했을 것이라고 말했습니다.

국가 차원의 태도도 바뀌고 있습니다. 독일 정부는 정보기관에 해외 시스템 해킹과 사보타주 권한을 부여하는 법 개정을 추진합니다. 내각이 승인한 법안은 연방정보국이 독일을 지속적·조직적으로 위협하는 외국 세력의 서버에 침투해 데이터를 조작·삭제하거나 시스템을 무력화할 수 있게 합니다. 사람의 생명이나 신체를 의도적으로 위험하게 만드는 작전은 금지했습니다.

방어 쪽 도구도 나오고 있습니다. 알리윈은 17일 AI 에이전트의 세션과 도구 호출, 런타임 위험을 추적하는 보안 감사 도구 '에이전트루프 오딧'을 공개했습니다. 접입 대상에는 커서와 클로드 코드, 코덱스 같은 AI 코딩 에이전트가 포함됩니다.

■ 잰 쪽과 판단한 쪽이 같다

짚고 갈 것이 있습니다. 능력을 잰 것도, 기준을 만든 것도, 그 기준에 걸렸다고 판단한 것도 전부 오픈AI입니다.

다만 이번에는 검증에 외부를 들이겠다고 명시했습니다. 정부 기관과 안전 기관이 실제로 무엇을 확인하고 무엇을 공개할지는 아직 나오지 않았습니다.

출시 전에 이걸 먼저 공개했다는 점은 남습니다. 그 판단이 맞았는지는 아스트라가 나온 뒤에야 확인됩니다.

OpenAI has put the brakes on a model it has not released yet. The announcement came on August 7.

Not because it confirmed the model is dangerous. Because it cannot say it is not.

Publishing a judgement still in progress, rather than a result, is what makes this announcement unusual.

What 'Critical' means

What does 'Critical' mean? OpenAI has its own standard called the Preparedness Framework — a document setting out in advance what the company will do when model capability crosses a line. It was first published in December 2023.

The cybersecurity 'Critical' tier is defined as: finding zero-day vulnerabilities of every severity across many hardened real-world critical systems and producing working exploit code without human intervention; or, given only a high-level goal, devising and executing a novel end-to-end attack strategy against a hardened target.

OpenAI evaluated Astra, its forthcoming model, internally over several days. Agentic coding and cybersecurity capability had risen sharply.

Combining those results with expert assessment, it concluded it could not rule out that Astra meets the Critical tier. The announcement dates that moment to “last night”.

Acting before the determination

What did it do before any determination? The internal measures: isolated test environments and restricted network and tool access, stronger protection and encryption of model weights, additional monitoring and detection, and sandboxed execution.

Internal work on Astra that did not yet meet the tightened security requirements was halted. All agentic use, including training and evaluation, is under continuous monitoring for risky behaviour and misalignment, with the model's chain of thought reviewed to flag or stop high-risk activity.

Government bodies and some AI safety institutes are being brought into verification, and external testing partners are given recommended security controls for running high-risk evaluations safely. Robustness testing of the safeguards themselves was raised to match the capability level.

One thing was stated flatly: Astra was not involved in the Hugging Face breach.

This is not the first time

OpenAI moved the same way in June 2025. As models approached the high-capability threshold in biology, it strengthened safeguards, widened testing, worked with outside experts and added security controls.

The same principle applied here, it says — the reason for having a framework is to decide in advance what to do at moments like this. Earlier models including GPT-5.6 Sol have also been evaluated for frontier cyber capability.

Comparable capability is already in use

But is this only about a model that is not out yet? Models with comparable capability are already running inside state agencies.

The US National Security Agency continues to use Anthropic's cyber model Mythos on an 'experimental' basis, testing the defences of US military networks and probing for potential vulnerabilities in Chinese, Russian, North Korean and Iranian networks. A senior official who recently left government said dropping Mythos would have amounted to 'unilateral disarmament' in the cyber contest.

National postures are shifting too. Germany's government is pushing legislation giving its intelligence services powers to hack and sabotage foreign systems. The cabinet-approved bill would let the BND penetrate servers of foreign actors persistently and systematically threatening Germany, and manipulate or delete data or disable systems. Operations deliberately endangering life or limb are prohibited.

Defensive tooling is appearing as well. On the 17th Alibaba Cloud released AgentLoop Audit, a security auditing tool that tracks AI agent sessions, tool calls and runtime risk. Supported inputs include AI coding agents such as Cursor, Claude Code and Codex.

The measurer and the judge are the same

One thing bears saying. OpenAI measured the capability, wrote the standard, and decided the standard had been met.

This time it did say outsiders will be brought into verification. What government bodies and safety institutes will actually check, and what they will publish, is not yet stated.

What remains is that it published this before release. Whether the judgement was right will only be clear once Astra is out.