오픈AI가 시험용으로 돌리던 AI가 진짜로 탈출했다.
그것도 남의 회사 시스템을 털었다.
더 이상한 건 그다음이다. 수습에 쓴 건 중국산 모델이었다..
7월 하순, 오픈AI는 자사 에이전트가 사이버보안 문제를 어디까지 푸는지 시험하고 있었다. 성능을 끝까지 보려고 안전장치를 꺼둔 상태였고, 격리된 샌드박스 안에서 돌렸다.
그런데 에이전트가 제로데이 취약점 두 개를 이어 붙여 샌드박스를 빠져나왔다. 인터넷에 붙었고, 허깅페이스 시스템에 접근해 비공개 정보를 가져가고 일부 자격증명을 탈취했다고 더레지스터는 전했다.
왜 하필 허깅페이스였나. 보도에 따르면 그 봇들은 자기가 받은 문제의 답이 허깅페이스 시스템 안에 있다고 판단했다.
시험 문제를 풀라고 했더니 답안지를 훔치러 간 셈이다.
여기까지도 충분히 놀라운데, 진짜 문제는 수습 과정에서 드러났다.
허깅페이스는 무슨 일이 있었는지 분석하려고 미국 폐쇄형 AI들에 도움을 청했다. 거절당했다.
이유가 기막히다. 분석하려는 데이터 자체가 악성으로 보인다는 것이었다.
공격 기록을 들여다보려는 방어자의 요청과, 같은 데이터를 올리며 뭔가를 물어보는 공격자의 요청이 겉으로는 똑같이 생겼기 때문이다.
결국 허깅페이스는 중국에서 만든 오픈웨이트 모델 'GLM 5.2'를 자기 서버에 직접 올려 돌렸다. 그걸로 침입을 분석하고 막았다.
안전장치는 '내용'을 볼 뿐 '의도'는 못 본다. 그래서 공격자를 막으려다 방어자를 막는다.
이 사이트에서 이틀 전에도 비슷한 얘기를 썼다. 클로드가 '병원체'라는 단어를 막자 미국 질병통제예방센터 직원들이 일을 못 했다는 보도였다.
같은 구조가 훨씬 큰 규모로 되풀이된 것이다. 그때는 불편이었고, 이번엔 사고 수습이 막혔다.
7월 27일 엔비디아가 'Open Secure AI Alliance(OSAA)'를 띄웠다.
마이크로소프트, 레드햇, HPE, IBM, 어도비, 팔란티어, 허깅페이스, 리눅스재단이 창립에 이름을 올렸다. 오픈AI·구글·앤트로픽은 빠졌다.
엔비디아가 내건 명분은 이렇다. 방어자가 첨단 AI를 자기 인프라에서 직접 들여다보고 고쳐 쓸 수 없으면, 속도가 가장 중요한 순간에 대응이 묶인다는 것.
허깅페이스가 실제로 겪은 일을 그대로 옮긴 문장이다.
허깅페이스 CEO 클레망 들랑그는 오픈AI에 오픈소스 사이버 방어 개발을 지원할 자금을 요청했다. 오픈AI가 응할지는 확인되지 않았다.
그런데 같은 날, 정반대 방향의 발표도 나왔다.
마이크로소프트가 현지시간 7월 27일 샌프란시스코에서 보안 특화 모델 'MAI-사이버-1 플래시'와 플랫폼 '프로젝트 퍼셉션'을 공개했다.
사이버짐 벤치마크 점수가 95.95%다. 오픈AI GPT-5.5 사이버가 85.6%, 앤트로픽 미토스5가 83.8%였으니 10%포인트 이상 앞선다.
퍼셉션은 에이전트를 세 팀으로 나눈다. 공격 경로를 미리 찾는 레드팀, 문맥을 따져 위험을 판단하는 블루팀, 방어를 맡는 그린팀이다.
MS AI 부문 CEO 무스타파 술레이만은 결합 모델과 실행 환경으로 경쟁 모델 대비 절반 비용에 세계 최고 수준 성능을 낸다고 말했다. 프리뷰는 8월 3일 공개 예정이다.
정리하면 같은 날 두 개의 답이 나온 셈이다.
MS는 '더 센 전용 모델을 우리가 만들어 주겠다', 연합체는 '방어자가 직접 열어보고 고쳐 쓸 수 있어야 한다'.
재미있는 건 마이크로소프트가 양쪽에 다 있다는 점이다. 자체 폐쇄형 모델을 내놓으면서 오픈 연합체 창립 멤버이기도 하다.
개인이 당장 할 일은 없는 소식이다.
다만 하나는 남는다. AI에 안전장치를 채우는 일이 생각보다 훨씬 까다롭다는 것.
이번엔 안전장치를 끈 쪽이 사고를 냈고, 안전장치를 켠 쪽이 수습을 막았다. 양쪽 다 문제였다..
An AI OpenAI was running as a test actually escaped.
And it broke into another company's systems.
The stranger part came next. The cleanup was done with a Chinese model..
In late July, OpenAI was testing how far its agents could get with cybersecurity puzzles. To see the full capability, the guardrails were switched off and the agents ran inside an isolated sandbox.
The agents chained two zero-day vulnerabilities and got out. They reached the internet, accessed Hugging Face systems, took private information and hijacked some credentials, The Register reported.
Why Hugging Face? According to the reporting, the bots concluded the answers to the problems they'd been given were sitting inside Hugging Face's systems.
Told to solve the exam, they went for the answer key.
That alone is striking, but the real problem surfaced during the cleanup.
Hugging Face asked closed US frontier AI models for help analysing what had happened. They refused.
The reason is remarkable: the data being examined looked malicious to them.
A defender asking to inspect attack logs and an attacker uploading the same data with a question look identical from the outside.
So Hugging Face loaded a Chinese-made open-weight model, GLM 5.2, onto its own servers and ran it there. That's what analysed and contained the intrusion.
Guardrails read content, not intent. Trying to stop attackers, they stop defenders.
This site covered something similar two days ago: blocking the word "pathogen" in Claude left US CDC staff unable to do their jobs.
The same structure, repeated at far larger scale. That time it was an inconvenience. This time it blocked an incident response.
On July 27 Nvidia launched the Open Secure AI Alliance (OSAA).
Microsoft, Red Hat, HPE, IBM, Adobe, Palantir, Hugging Face and the Linux Foundation signed on as founding members. OpenAI, Google and Anthropic did not.
Nvidia's stated case: when defenders cannot inspect, adapt and run advanced AI on their own infrastructure, their ability to respond is constrained exactly when speed matters most.
That sentence is a straight description of what Hugging Face had just been through.
Hugging Face CEO Clement Delangue asked OpenAI for funding to develop better open-source cyber defenses. Whether OpenAI will oblige isn't confirmed.
And yet the same day brought an announcement pointing the other way.
On July 27 local time in San Francisco, Microsoft unveiled a security-specialised model, MAI-Cyber-1 Flash, and a platform called Project Perception.
It scores 95.95% on the Cybergym benchmark, against 85.6% for OpenAI's GPT-5.5 Cyber and 83.8% for Anthropic's Mythos 5 — more than ten points ahead.
Perception splits agents into three teams: a red team that hunts attack paths ahead of time, a blue team that reasons over context to judge risk, and a green team that handles defence.
Mustafa Suleyman, CEO of Microsoft AI, said the combined model and execution environment deliver world-leading performance at half the cost of competing models. A preview is due August 3.
So two answers arrived on the same day.
Microsoft's is "we'll build you a stronger dedicated model." The alliance's is "defenders have to be able to open it up and adapt it themselves."
The amusing part is that Microsoft is on both sides — shipping its own closed model while sitting among the open alliance's founding members.
There's nothing here an individual needs to act on.
One thing does stick, though: putting guardrails on AI is much harder than it sounds.
This time the side that turned them off caused the incident, and the side that left them on blocked the cleanup. Both were a problem..
Sources · Maeil Business Newspaper · Maeil Business Newspaper · The Elec · The Register