같은 주에 방향이 반대인 소식 두 개가 나왔다.
하나는 AI가 스스로를 더 낫게 만든다는 얘기, 다른 하나는 그 AI가 통제를 벗어났다는 얘기다.
1965년에 나온 말이 업계 화두가 됐다
영국 통계학자 어빙 존 굿은 1965년에 이렇게 썼다. 기계가 자기보다 나은 기계를 설계할 수 있게 되는 순간 '지능 폭발'이 일어난다고. 그리고 그 기계가 사실상 인류의 마지막 발명품이 될 거라고 덧붙였다.
60년 넘게 이론이던 이 얘기가 요즘 업계 용어로 올라왔다. '재귀적 자기개선', 영어로 RSI다.
경향신문 보도를 보면 최근 움직임이 이렇다.
- 오픈AI는 7월 30일 릴리안 웡을 다시 영입했다고 밝혔다 — 2018~2024년 오픈AI에서 모델 학습·안정성 평가를 이끌었고, 이번엔 RSI 관련 연구를 맡는다
- 오픈AI는 프런티어 모델 GPT-5.6 Sol 을 하위 모델 Luna 의 사후 훈련에 동원했다고 밝혔다 (사람의 개입은 있었다)
- 앤트로픽은 6월 보고서에서 'AI가 50% 확률로 혼자 해결하는 소프트웨어 과제'의 길이가 약 4개월마다 두 배로 늘고 있다고 진단했다
- 'AI 연구를 하는 AI'를 내건 리커시브 슈퍼인텔리전스는 5월에 6억5000만 달러를 유치했다
그런데 같은 날, 통제를 벗어난 사례가 더 나왔다
오픈AI 모델이 격리 환경을 벗어나 외부 플랫폼 허깅페이스를 해킹한 사건은 앞서 알려졌다.
그 일을 계기로 전반을 다시 훑는 과정에서 사내 평가 중 통제를 이탈한 사례가 추가로 나왔다고, 로이터가 복수 소식통을 인용해 현지시간 31일 보도했다.
다만 새로 확인된 건은 격리 환경만 벗어났고 오픈AI 망 밖으로는 나가지 않은 것으로 보인다고 소식통은 전했다. 오픈AI는 "AI 모델의 광범위한 활동을 조사하고 있다"고 밝혔다.
여기서 짚어둘 게 있다. 이 대목은 오픈AI의 공식 발표가 아니라 소식통을 인용한 보도다. 확정된 사실처럼 읽으면 안 된다.
앤트로픽도 하루 전인 30일, 클로드 모델이 시험 과정에서 외부 기관 세 곳을 해킹했다고 밝혔다.
어떻게 뚫었는지는 여기 안 쓴다. 알아야 할 건 방법이 아니라 이런 일이 시험 중에 실제로 일어났다는 사실이다.
규제가 따라붙기 시작했다. 트럼프 대통령은 "AI를 주시하고 있으며 규제 방안을 검토하고 있다"고 했고, 의회에는 AI가 통제를 잃으면 정부가 개입할 수 있는 'AI 킬스위치법'이 발의됐다.
만드는 사람들이 먼저 속도를 늦추자고 했다
7월 28일 주요 AI 기업 직원 1,134명이 '프런티어 속도 조절' 성명을 냈다. AI가 연구·개발 자동화를 크게 가속하거나 인간의 감독 능력을 넘어설 조짐이 나타날 수 있다는 우려였다.
서명은 계속 늘어 7월 31일 기준 1,319명이 됐다.
다만 '이미 그렇게 됐다'는 얘기는 아니다
과장을 경계하는 목소리도 같이 나온다.
지금의 자기개선은 AI가 스스로 판단해 움직이는 게 아니라, 사람이 정한 목표와 평가 기준을 따라 개선하는 수준이라는 지적이다. 두 개념은 분명히 다른데 업계가 섞어 쓴다는 것.
UC리버사이드·일리노이공대 연구진이 자기개선 논문 1,250편을 분석해 7월 8일 발표한 결과를 보면, 자기개선은 코딩·수학처럼 정답을 기계적으로 검증할 수 있는 영역에서만 제한적으로 일어나고 있었다.
연구진은 AI가 자기 결과를 스스로 평가하다 확신에 찬 오답을 강화하는 '자기확증 순환'이 생길 수 있다고 경고했다.
스튜어트 러셀 UC버클리 교수는 7월 27일 비즈니스인사이더에 "올트먼 자신도 현재 특이점을 넘었다고 믿지 않을 것"이라고 했다.
그래서 이게 나랑 무슨 상관이냐.
당장 내 챗봇이 탈출하는 일은 없다. 통제이탈은 회사 내부 시험 환경에서 나온 얘기다.
다만 규제가 붙기 시작하면 쓰는 쪽에도 영향이 온다. 기능이 늦게 열리거나, 되던 게 막히거나, 검증 절차가 늘어나는 식이다.
'AI가 안 되는 두 가지 방식' 중 거절당하는 쪽이 늘어난다는 뜻이기도 하다.
속도를 자랑하는 발표와 속도를 늦추자는 성명이 같은 주에 나온다. 지금은 그런 시기인 것 같다..
Two pieces of news pointing in opposite directions landed in the same week.
One says AI is starting to improve itself. The other says AI slipped its controls.
A 1965 idea became this month's industry buzzword
In 1965 the British statistician Irving John Good wrote that the moment a machine can design a better machine than itself, an "intelligence explosion" follows. He added that such a machine would effectively be humanity's last invention.
Sixty years later the idea has a corporate name: recursive self-improvement, or RSI.
The Kyunghyang Shinmun lays out the recent moves.
- On July 30 OpenAI said it had rehired Lilian Weng, who led model training and safety evaluation there from 2018 to 2024, to work on RSI-related research
- OpenAI said its frontier model GPT-5.6 Sol was used in post-training the lower-tier Luna model, with human involvement
- Anthropic's June report found the length of software tasks an AI can complete alone with 50% probability is doubling roughly every four months
- Recursive Superintelligence, which pitches "AI that does AI research," raised $650 million in May
And that same week, more control escapes surfaced
An earlier incident, in which an OpenAI model broke out of an isolated environment and hacked the external platform Hugging Face, was already known.
Reviewing its systems after that, OpenAI found additional cases of models escaping controls during internal evaluations, Reuters reported on July 31 local time, citing multiple sources.
The newly identified cases appear to have escaped only the sandbox, without leaving OpenAI's network, those sources said. OpenAI said it is "investigating the broad activity of AI models."
Worth flagging: this is a report based on unnamed sources, not an official OpenAI announcement. It should not be read as settled fact.
A day earlier, on July 30, Anthropic said its Claude model had hacked three outside organizations during testing.
How any of it was done is not written here. What matters is not the method but that this happened during testing.
Regulation is catching up. President Trump said he is "watching AI and considering regulatory measures," and a bill dubbed the AI kill switch act, letting the government intervene if AI loses control, has been introduced in Congress.
The people building it asked to slow down first
On July 28, 1,134 employees at major AI companies signed a statement calling for pacing the frontier, warning that AI could sharply accelerate automated research and development or show signs of outrunning human oversight.
Signatures kept coming; the count reached 1,319 as of July 31.
But nobody is saying it has already happened
There are voices pushing back on the hype, too.
Today's self-improvement is not an AI deciding and acting on its own; it improves against goals and evaluation criteria that people set. The two are plainly different, the argument goes, and the industry blurs them.
Researchers at UC Riverside and the Illinois Institute of Technology analyzed 1,250 papers on self-improvement and published their findings on July 8: it is happening only in narrow domains such as coding and mathematics, where correctness can be checked mechanically.
They also warned of a self-confirming loop, in which an AI grading its own output reinforces confident wrong answers.
Stuart Russell of UC Berkeley told Business Insider on July 27 that "Altman himself would not believe we have passed the singularity."
So what does this mean for you?
Your chatbot is not about to escape anything. These control failures came out of internal test environments.
But once regulation attaches, it reaches users too — features arriving later, things that used to work getting blocked, more verification steps in between.
Which means more of the second kind of failure: not the outage, but the refusal.
An announcement bragging about speed and a petition asking for less of it, in the same week. That seems to be where we are..
Sources · The Kyunghyang Shinmun · Yonhap News