앤트로픽이 AI 가 다른 AI 모델의 정렬 성능을 개선할 수 있음을 보여주는 연구 결과를 내놨습니다.

28일 공개한 논문 '자동화된 연구자는 정렬 실패를 효과적으로 완화할 수 있다'입니다. 펠로우 프로그램 연구원 천 유에한이 주도했습니다.

정렬은 AI 가 인간의 의도와 가치에 맞게 행동하도록 만드는 기술입니다. 그 일을 AI 에게 시켜본 실험입니다.

■ 10개 벤치마크 모두에서

이 시스템의 이름은 '자동화된 AI 정렬 연구자(AAR)'입니다. AI 가 연구자 역할을 맡아 다른 모델의 정렬을 개선합니다.

결과는 어땠을까요? 논문에 따르면 클로드 기반 자동화 시스템은 10개 오정렬 벤치마크 모두에서 소형 AI 모델의 정렬 성능을 개선했고, 전체 모델 성능 저하도 발생하지 않았습니다.

■ 사람의 방법을 그대로 흉내 낸다

어떻게 그게 가능할까요? 작동 방식은 특별하지 않습니다. 인간 연구자의 방법론을 그대로 모방합니다.

각 자동화 시스템이 기존 문헌을 검색하고, 방법론을 제안하고, 30분간 모델을 훈련시킵니다. 이를 여러 차례 반복하며 효과적인 방법은 남기고 그렇지 않은 것은 버립니다. 다른 점은 속도와 규모입니다.

논문에 따르면 최고 성능의 AAR 은 평균 6시간 만에 숙련된 인간 연구자의 제안을 능가했습니다. 인간이 제시한 연구 방향도 성능 향상에 별다른 도움이 되지 않은 것으로 나타났습니다.

■ 시간당 4달러와 150달러

격차가 더 큰 쪽은 비용입니다.

AAR 의 API 추론 비용은 시간당 약 4달러인 반면 인간 연구자는 150달러 수준으로, 약 37배 저렴한 것으로 나타났습니다.

AI 해설가 로한 폴은 28일 엑스에서 이번 연구를 '재귀적 자기개선' AI 의 중요한 진전이라고 평가했습니다. 클로드가 48시간 동안 스스로 안전성 개선 방안을 찾아 인간 연구자 28명의 아이디어를 뛰어넘었으며, AI 가 더 강력한 후속 AI 를 안전하게 만드는 연구까지 수행할 수 있는 단계에 다가서고 있다는 것입니다.

■ 자동화가 만능은 아니다

그러면 사람은 이제 필요 없을까요? 연구진은 선을 그었습니다. AAR 의 가능성을 인정하면서도, 벤치마크와 참고 문헌의 지속적인 관리가 필요해 자동화가 만능은 아니라고 강조했습니다.

결국 무엇을 기준으로 평가할지, 무엇을 읽힐지는 사람이 정해야 한다는 뜻입니다. 자동화된 것은 그 사이의 반복 작업입니다.

■ 같은 주에 법원에서

이 발표가 나온 자리도 짚어둘 만합니다. 앤트로픽은 같은 시기 미국 국방부와의 소송에서 이겼습니다.

캘리포니아 연방지방법원 리타 린 판사는 27일 국방부의 앤트로픽 '공급망 위험' 지정이 위법이라고 판결했습니다. 올해 2월 헤그세스 국방장관이 자율무기와 국민 대규모 감시에 AI 사용을 거부한 앤트로픽을 공급망 위험으로 지정해 군 전체의 클로드 사용을 금지한 조치였습니다.

법원은 “국가안보의 공허한 원용은 정부 비판자를 처벌하는 백지 수표가 아니다”라며 제1수정헌법 위반으로 판시했습니다.

AI 가 AI 를 고치는 실험, 그리고 군에서 쫓겨났다가 법정에서 돌아온 회사. 같은 주에 나란히 놓인 소식입니다.

Anthropic has published research showing that AI can improve the alignment of another AI model.

The paper, 'Automated researchers can effectively mitigate alignment failures', was released on the 28th and led by Chen Yuehan, a researcher in the company's fellowship programme.

Alignment is the work of making AI behave in line with human intent and values. This experiment handed that work to AI.

All ten benchmarks

The system is called an automated AI alignment researcher, or AAR: AI takes the researcher's role and improves alignment in another model.

How did it do? According to the paper, the Claude-based system improved alignment in a small AI model on all ten misalignment benchmarks, with no degradation in overall model performance.

It copies the human method

How is that possible? There is nothing exotic about how it works. It imitates a human researcher's methodology.

Each system searches the existing literature, proposes a method, and trains a model for 30 minutes. It repeats this many times, keeping what works and discarding what does not. What differs is speed and scale.

The best-performing AAR surpassed proposals from experienced human researchers in six hours on average. Research directions suggested by humans turned out to add little to performance.

$4 an hour against $150

The larger gap is cost.

AAR inference through the API costs about $4 an hour, against roughly $150 for a human researcher — about 37 times cheaper.

The commentator Rohan Paul called it an important step toward recursive self-improvement on X on the 28th, noting that Claude spent 48 hours finding safety improvements on its own and surpassed the ideas of 28 human researchers, moving closer to AI that can do the research to make a more powerful successor safe.

Automation is not everything

So are humans no longer needed? The researchers drew a line. They acknowledged AAR's potential while stressing that benchmarks and reference literature need continuous maintenance, so automation is not a cure-all.

In other words, people still have to decide what to evaluate against and what to read. What has been automated is the repetition in between.

In court the same week

The setting of this announcement is worth noting. In the same period, Anthropic won its case against the US Department of Defense.

On the 27th, Judge Rita Lin of the US District Court in California ruled the Pentagon's designation of Anthropic as a 'supply chain risk' unlawful. In February, Defense Secretary Hegseth had applied that designation to Anthropic — which refused to allow its AI in autonomous weapons and mass domestic surveillance — barring Claude across the military.

The court held it violated the First Amendment, saying that a hollow invocation of national security is not a blank cheque to punish the government's critics.

An experiment in AI repairing AI, and a company thrown out of the military that came back through a courtroom. Two items from the same week.

Sources · Choice Economy