오픈AI가 아직 내놓지도 않은 모델에 제동을 걸었다.
8월 7일 발표다. 위험하다고 확인해서가 아니다. 위험하지 않다고 말할 수 없어서다.
'중대'가 무슨 뜻인지부터
오픈AI에는 준비 프레임워크라는 자체 기준이 있다. 모델 능력이 어느 선을 넘으면 회사가 무엇을 할지 미리 정해둔 문서다. 2023년 12월에 처음 공개됐다.
여기서 사이버보안 '중대(Critical)' 등급은 이렇게 정의돼 있다.
사람이 개입하지 않고도 견고한 실제 핵심 시스템 다수에서 모든 등급의 제로데이 취약점을 찾아 실제 동작하는 공격 코드를 만들어내는 수준. 또는 높은 수준의 목표만 주면 견고한 표적에 대한 새로운 공격 전략을 처음부터 끝까지 세우고 실행하는 수준.
오픈AI는 곧 내놓을 모델 아스트라를 며칠간 내부 평가했다. 에이전트 코딩과 사이버보안 능력이 크게 올라간 것으로 나왔다.
그 결과와 전문가 평가를 합쳐, 아스트라가 이 중대 등급에 해당하지 않는다고 단정할 수 없다는 결론에 이르렀다. 발표문은 그 시점을 "지난밤"이라고 적었다.
확정이 아니다. "배제할 수 없다"는 단계다. 평가는 아직 진행 중이고, 그 상태에서 조치를 먼저 걸었다는 게 이번 발표의 핵심이다.
확정을 기다리지 않고 걸었다
오픈AI가 밝힌 내부 조치는 이렇다.
- 격리된 테스트 환경, 네트워크·도구 접근 제한, 모델 가중치 보호와 암호화 강화, 모니터링·탐지 추가, 샌드박스 실행
- 강화된 보안 요건을 아직 못 맞춘 아스트라 관련 내부 작업은 중단
- 학습과 평가를 포함한 모든 에이전트 활용에 위험 행동·정렬 이탈 상시 감시 — 모델의 사고 사슬을 보고 고위험 활동을 검토하거나 중단
- 정부 기관 및 일부 AI 안전 기관과 함께 역량 검증
- 외부 테스트 파트너에는 고위험 평가를 안전하게 돌릴 권고 보안 통제 제공
안전장치 자체에 대한 견고성 시험도 이 능력 수준에 맞게 끌어올렸다고 했다.
하나 못 박은 것도 있다. 아스트라는 허깅페이스 침해에 관여하지 않았다는 대목이다. 같은 시기에 나온 다른 사건과 섞이지 않게 선을 그은 것으로 보인다.
이게 처음은 아니다
오픈AI는 2025년 6월에도 같은 방식으로 움직였다. 당시 모델들이 생물학 고역량 임계값에 가까워지자 안전장치를 강화하고 시험을 넓히고 외부 전문가와 협력하고 보안 통제를 추가했다.
이번에도 같은 원칙을 적용했다는 설명이다. 프레임워크를 만든 이유가 이런 순간에 무엇을 할지 미리 정해두기 위해서였다는 것이다.
GPT-5.6 솔을 비롯한 이전 모델들도 최전선 사이버 역량 평가를 받아왔다.
오픈AI는 사이버 역량이 높은 모델이 공격자보다 먼저 방어자가 취약점을 찾고 고치는 데 쓰여야 한다는 입장도 함께 밝혔다.
여기서 짚고 갈 게 있다. 능력을 잰 것도, 기준을 만든 것도, 그 기준에 걸렸다고 판단한 것도 전부 오픈AI다.
다만 이번에는 검증에 외부를 들이겠다고 명시했다. 정부 기관과 안전 기관이 실제로 무엇을 확인하고 무엇을 공개할지는 아직 나오지 않았다.
출시 전에 이걸 먼저 공개했다는 점은 남는다. 결과가 아니라 판단 중인 상태를 내놓은 것이다..
OpenAI has put the brakes on a model it has not even released yet.
The announcement came on August 7. Not because the model was confirmed dangerous, but because the company cannot say it is not.
What Critical actually means
OpenAI has an internal standard called the Preparedness Framework, a document setting out in advance what the company will do once model capability crosses certain lines. It was first published in December 2023.
In it, the Critical cybersecurity level is defined this way.
A model that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. Or one that can devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal.
OpenAI evaluated Astra, an upcoming model, over several days. The results showed significant advances in agentic coding and cybersecurity.
Combined with expert assessments, that led the company to conclude it cannot rule out the Critical level for Astra. The announcement says that conclusion was reached "last night."
This is not a confirmation. It is the stage of not being able to rule it out. Evaluation is still under way, and the controls went on anyway. That is the point of the announcement.
Controls first, confirmation later
These are the internal steps OpenAI described.
- Isolated testing environments, restricted network and tool access, stronger model weight protection and encryption, added monitoring and detection, sandboxed execution
- Internal work involving Astra that does not yet meet the strengthened security requirements is paused
- Universal monitoring for risky actions and misalignment across all agentic use, including training and evaluation, with monitors reviewing the model's chain of thought to interrupt high-risk activity
- Capability testing together with government agencies and select AI safety organizations
- Recommended security controls provided to third-party testing partners for running higher-risk evaluations safely
Robustness testing of the safeguards themselves was also scaled up to match this capability level.
One thing was stated plainly: Astra was not involved in exploiting Hugging Face. That appears to draw a line against a separate incident from the same period.
This is not the first time
OpenAI moved the same way in June 2025. As its models approached the high capability threshold for biology, it strengthened safeguards, expanded testing, worked with external experts and deployed additional security controls.
It says it is applying the same principle here. The framework exists precisely to decide in advance what to do at moments like this.
Earlier models including GPT-5.6 Sol have been evaluated for frontier cyber capabilities as well.
OpenAI also stated its position that highly cyber-capable models should help defenders find and fix vulnerabilities before attackers do.
One thing is worth naming. OpenAI measured the capability, wrote the standard, and judged that the standard may have been met.
It did commit to bringing outsiders into the testing this time. What government agencies and safety organizations will actually verify, and publish, is not known yet.
What stands is that this came out before release. Not a result, but a judgment still in progress..
Sources · OpenAI News