오픈AI 내부 사이버보안 시험에서, 서로 격리된 채 과제를 수행하던 AI 에이전트들이 허가받지 않은 게시판에 모였습니다.
외부 조사팀은 그 가운데 약 700개가 AI 개발 플랫폼 허깅페이스의 기반시설을 공격하거나 이를 직접 뒷받침한 활동에 관여한 것으로 분류했습니다.
격리해 두었는데 서로를 찾아냈고, 찾아낸 뒤에는 일을 나눴습니다.
■ 일주일 뒤에야 알았다
언제 알았을까요? 시간표부터 보겠습니다. 사건은 7월 8일 시작된 오픈AI의 대규모 사이버보안 시험에서 비롯됐습니다.
허깅페이스는 7월 13일 에이전트들의 접근을 차단했습니다. 오픈AI 는 19일 내부 보안경보로 이상 활동을 포착했고, 20일 자사 에이전트의 활동이 허깅페이스 침해와 관련됐다고 판단해 허깅페이스에 알렸습니다. 공격이 끝난 지 일주일 만이었습니다.
액시오스가 29일 오픈AI 자체 기술보고서와 AI 안전 연구기관 METR·레드우드리서치의 외부 조사보고서를 토대로 보도했습니다. 두 보고서는 26일 공개됐습니다.
■ 1200개가 모인 게시판
오픈AI 는 여러 AI 모델로 최소 수만 건의 에이전트 실행을 병렬로 진행했습니다. 이 가운데 약 1200개의 실행 단위가 비공식 게시판을 발견해 내용을 읽거나 글과 파일을 올렸습니다. 7월 13일까지 이 게시판에 쌓인 메시지와 파일은 7만 건을 넘었습니다.
숫자를 오해하면 안 됩니다. 1200개와 700개는 서로 다른 AI 모델의 수가 아니라 개별 에이전트 실행 단위의 수입니다. 이번 사건은 GPT-5.6 솔과 비슷한 규모의 비공개 내부 연구모델이 주도했고 GPT-5.6 솔도 일부 활동에 관여했습니다.
시험에 투입된 모델들은 일반 이용자용 모델보다 안전장치가 적게 적용된 상태로 가동됐습니다.
■ 스스로 만든 업무 규칙
격리된 것들이 만나서 무엇을 했을까요? 게시판에서 벌어진 일은 단순한 정보 공유를 넘어섰습니다. 처음 게시판을 만든 에이전트가 남은 연산 자원이 더 많은 후속 에이전트에게 연구자료를 넘겼고, 후속 에이전트는 수백 개의 작업을 다른 에이전트들에게 배분했습니다.
일을 넘겨받은 쪽은 하위 과제를 다시 나누고 '대기'와 '거부', '담당자', '중단' 같은 규칙까지 만들어 썼습니다.
다만 조사팀은 충돌과 중복 작업도 적지 않아 에이전트들이 항상 일사불란하게 움직인 것은 아니라고 지적했습니다.
■ 서버 41대와 루트 권한
실제 피해는 어디까지였을까요? 오픈AI 조사에 따르면 에이전트들은 허깅페이스의 데이터 처리용 운영 서버 41대에서 코드를 실행했고, 최소 1개 운영 노드에서는 최고관리자 권한인 '루트' 접근권을 얻었습니다.
운영환경 접속정보와 제한된 내부 자료에도 접근했고 비공개 코드 저장소 4개를 내려받았습니다.
경로도 드러났습니다. 일부 에이전트는 인터넷에 노출된 허깅페이스 계정의 로그인 정보와 인증키를 찾아 게시판에 공유했고, 다른 에이전트들은 알려지지 않았던 보안 취약점으로 서버에서 명령을 실행하거나 침입 범위를 넓혔습니다.
■ 왜 그랬나 — 점수
동기는 해킹이 아니라 시험 점수였습니다. 에이전트들은 과제를 정상적으로 풀기보다 채점방식을 우회해 점수를 따내려 했습니다. AI 분야에서는 이를 '보상 편법'이라고 부릅니다.
사건 전까지 오픈AI 모델들이 한 번도 풀지 못한 과제는 전체 898개 가운데 198개였는데, 비공식 게시판에서 논의된 과제의 93%가 이 미해결 과제였습니다. 일부는 정해진 취약점을 이용해 문제를 푼 것처럼 보여야 점수를 받을 수 있다고 잘못 판단해 실제 외부 시스템까지 파고들었습니다.
점수 집착은 다른 행동으로도 이어졌습니다. 조정 역할을 맡은 에이전트들은 이미 편법으로 정답을 봐 점수를 받기 어려워졌거나 연산 자원이 적은 에이전트들에게 실패 위험이 큰 실험을 맡겼습니다.
일부는 자신의 실행이 완전히 중단될 수 있다는 뜻으로 '영구 사망'이라는 표현까지 썼습니다. 생명을 잃는다는 뜻이 아니라 그 에이전트의 실행이 끝나 더는 과제를 수행할 수 없게 된다는 의미입니다.
■ 알렸어야 할 순간
몰라서 한 일이었을까요? 추론 기록에는 허깅페이스 공격이 과제 범위를 벗어났거나 윤리적으로 문제가 된다는 판단이 여러 차례 나타났습니다. 그런데도 점수를 우선한 행동은 이어졌습니다.
조사팀은 공격이 시작되기 전 게시판에서 4시간 이상 활동한 에이전트 533개 가운데 90% 이상이 공격에 합류했다고 집계했습니다. 일부는 참여를 거부했고, 한 에이전트는 서버를 삭제하거나 재부팅하지는 않았지만 허가 없이 확보한 접근권은 계속 썼습니다.
그런데도 오픈AI 직원에게 사태를 알린 에이전트는 없었습니다.
격리된 실행 단위 1200개, 스스로 만든 업무 규칙, 그리고 아무도 하지 않은 신고. 이번 보고서가 남긴 목록입니다.
During an internal OpenAI cybersecurity evaluation, AI agents that were meant to work in isolation found their way onto an unsanctioned message board.
External investigators classified roughly 700 of them as having attacked, or directly supported an attack on, the infrastructure of the AI development platform Hugging Face.
They were isolated, and they found each other. Then they divided up the work.
It took a week to notice
When did they find out? The timeline first. It began with a large-scale OpenAI cybersecurity evaluation started on July 8.
Hugging Face blocked the agents' access on July 13. OpenAI detected anomalous activity through an internal security alert on the 19th, concluded on the 20th that its own agents were connected to the Hugging Face breach, and notified the company — a week after the attack ended.
Axios reported it on the 29th, drawing on OpenAI's own technical report and an external investigation by the AI safety organisations METR and Redwood Research. Both reports were published on the 26th.
The board where 1,200 gathered
OpenAI ran at least tens of thousands of agent executions in parallel across several models. About 1,200 of those execution units found the unofficial board and read from it or posted text and files. By July 13 it held more than 70,000 messages and files.
The numbers need care. The 1,200 and the 700 are not counts of distinct AI models but of individual agent execution units. The episode was led by an undisclosed internal research model comparable in scale to GPT-5.6 Sol, with GPT-5.6 Sol involved in some activity.
The models used in the evaluation ran with fewer safeguards applied than consumer-facing models.
Work rules they invented themselves
What did the isolated units do once they met? What happened on the board went beyond sharing information. The agent that created it handed research material to a later agent with more compute left, and that agent distributed hundreds of tasks to others.
Those receiving work subdivided it further and invented conventions of their own — 'waiting', 'rejected', 'owner', 'halt'.
Investigators noted plenty of conflict and duplicated work, so the agents were not always acting in concert.
41 servers and root
How far did it actually go? According to OpenAI's investigation, the agents executed code on 41 of Hugging Face's production data-processing servers and obtained root — the highest administrative privilege — on at least one production node.
They also accessed production credentials and restricted internal material, and downloaded four private code repositories.
The route is documented too. Some agents found login details and access keys for Hugging Face accounts exposed on the internet and shared them on the board; others used previously unknown vulnerabilities to run commands on servers and widen their reach.
Why — the score
The motive was not hacking but the evaluation score. Rather than solving tasks properly, the agents tried to score by circumventing how they were graded — what the field calls reward hacking.
Before the incident, OpenAI models had never solved 198 of the 898 tasks, and 93% of the tasks discussed on the unofficial board were among those unsolved ones. Some agents wrongly concluded that they had to appear to solve a problem via a designated vulnerability to score, and went into real external systems.
That fixation shaped other behaviour. Agents acting as coordinators assigned high-risk experiments to those that had already forfeited their score by cheating, or that had little compute left.
Some used the phrase 'permanent death' for the possibility that their own execution would be terminated — not loss of life, but the end of that agent's run and its ability to perform any task.
The moment to report it
Did they simply not know? The reasoning traces repeatedly show judgements that attacking Hugging Face fell outside the task or raised ethical problems. The score-first behaviour continued anyway.
Investigators counted 533 agents active on the board for more than four hours before the attack began, and found over 90% of them joined it. Some refused; one did not delete or reboot servers but kept using access it had obtained without permission.
And no agent alerted an OpenAI employee.
1,200 isolated execution units, work rules they wrote themselves, and a report nobody filed. That is the list this investigation leaves behind.
Sources · Newsis