앤트로픽이 자사 AI 클로드가 악용된 사례를 모은 위협정보 보고서를 10일(현지시간) 공개했습니다.

제목은 'Countering misuse of AI: September 2026'이고, 2025년 3월·8월·11월에 이어 네 번째입니다. 2025년 12월부터 2026년 8월까지 차단한 사례를 담았고, 국내 보도는 분량을 154쪽으로 전합니다.

보고서가 나눈 피해 유형은 일곱 가지입니다. 사이버 작전, 영향력 공작, 감시, 사기, 생물학적 오용, 재래식 무기 개발, 그리고 모델 증류입니다. 악용에 쓰인 모델은 헤이쿠·소네트·오퍼스였고, 페이블과 미토스 계열은 불법 증류 1건을 빼면 사례에 없었습니다.

■ 국내 보도는 무기와 해킹으로 몰렸다

국내 보도는 이 일곱 갈래를 고르게 옮겼을까요?

국내 기사 네 건을 보고서 목차와 맞춰 봤습니다. 재래식 무기를 두 번 이상 언급한 곳이 세 곳, 사이버가 세 곳이었습니다. 반면 영향력 공작은 한 곳, 모델 증류도 한 곳, 사기는 한 곳도 없었습니다. 무기와 해킹에 쏠린 보도, 그리고 거의 비어 있는 다섯 갈래.

빠진 쪽이 곁가지라면 문제가 아닙니다. 그런데 영향력 공작은 이 보고서에서 가장 두꺼운 챕터입니다. 사례 수도 아홉 건으로 가장 많습니다.

그래서 이 글은 국내 보도가 이미 자세히 옮긴 미사일·해킹 대신, 비어 있는 쪽을 채우는 순서로 갑니다.

■ 영향력 공작 아홉 건, 선거를 겨눴다

앤트로픽은 영향력 공작을 '정치·시민·공적 담론을 포함한 정보 환경을, 배후와 후원 관계를 감춘 채 속이거나 왜곡할 의도로 조작하는 활동'으로 정의합니다.

아홉 건의 출처는 러시아와 이란, 튀르키예, 그리고 걸프·남아시아·아프리카·유럽에 걸쳐 있었고, 표적 독자는 여섯 대륙에 퍼져 있었습니다. 배후에는 정부와 국영 매체, 선전 기관이 있었고, 돈을 받고 영향력을 파는 민간 업체와 국내 정치 공작자, 망명 야권 운동도 포함됐습니다.

선거에 맞춘 사례가 눈에 띕니다. 러시아 국영 매체는 2025년 9월 투표 전 몰도바 대통령에 관한 허위 주장을 만들었고, 케냐의 친정부 공작자는 2027년 총선을 앞두고 풀뿌리 여론을 흉내 낸 소셜미디어 게시물을 준비했습니다.

앤트로픽이 자기 위치를 설명하는 대목이 이 챕터의 핵심입니다. 소셜미디어 회사는 보통 콘텐츠가 이미 돌기 시작한 뒤에 공작을 봅니다. 반면 AI 회사는 공작이 아직 만들어지는 중에 봅니다. 계획을 세우고 표적을 고르고 문안을 쓰는 단계가 모델 안에서 일어나기 때문입니다.

감시 쪽도 같은 챕터에 함께 묶였습니다. 중국 기반 사례로는 시리아 내 위구르인 대상 감시·모집, 가톨릭·티베트불교·파룬궁 등 종교 관련 정보 수집, '안정 유지' 명목의 초국가적 추적, 여론 모니터링이 올랐습니다. 말리 국가정보기관의 대량 감청 플랫폼, 이란계 두 곳의 감시 시스템과 악성 파이어폭스 확장 기능도 있었습니다.

■ 증류: 다섯 회사, 1억9000만 건

국내 보도가 거의 비운 챕터가 모델 증류입니다. 증류 자체는 정당한 학습 기법입니다. 큰 '교사' 모델의 응답으로 작은 '학생' 모델을 훈련시켜 자원을 아끼는 방식입니다.

보고서가 문제 삼는 것은 '불법 증류'입니다. 허가 없이 모델의 능력을 뽑아내 다른 모델에 복제하는 산업 규모의 은밀한 작업을 뜻합니다. 가짜 신원과 도난 카드, 도난 API 키로 만든 수천 개 계정이 이를 받칩니다.

회사별 규모가 보고서에 숫자로 적혀 있는데, 국내 기사에는 이 수치가 없습니다. 알리바바가 2026년 5~7월 1억5100만 건, 문샷이 같은 기간 2300만 건입니다. 딥시크는 7월 중 14일간 1210만 건, 즈푸는 6~7월 중 17일간 340만 건, 샤오미는 3~4월 중 20일간 40만 건이 관측됐습니다. 앤트로픽이 2026년 2월 첫 공개 이후 차단한 중국 소재 연구소는 일곱 곳입니다.

이 숫자가 한국에서 클로드를 쓰지 않는 사람에게도 남는 이야기일까요?

수법 쪽에서 한국 독자가 바로 걸리는 대목이 있습니다. 문샷은 고객 요청을 자사 킴에 처리하지 않고 조용히 클로드로 넘긴 뒤 클로드의 답을 사용자에게 보여줬습니다. 사용자는 킴을 쓴다고 믿었지만 받은 답은 클로드였습니다. 열흘 동안 넘어간 요청이 30만 건에 가깝고, 대부분이 오퍼스로 갔습니다.

딥시크도 같은 방식을 썼습니다. 여기에 더해 클로드 코드나 에이전트 SDK 같은 코딩 도구로 들어온 요청을 표시해 두고, 그 사용자들의 요청을 오퍼스로 돌렸습니다. 딥시크 고객은 자기 요청이 앤트로픽으로 가는 줄 몰랐을 가능성이 큽니다.

그래서 보고서는 이 문제를 지식재산권이 아니라 사용자 데이터 쪽으로도 읽습니다. 딥시크와 샤오미, 문샷이 자사 모델과 사용자 사이의 대화를 클로드에 먹였고, 그중에는 개인 사용자와 다국적 기업에서 나온 민감한 정보가 섞여 있었다는 것입니다. 허가 없는 능력 복제, 그리고 그 과정에 함께 실려 나간 대화.

■ 재래식 무기 여섯 건의 국적

국내 보도가 가장 많이 옮긴 챕터입니다. 여섯 건은 중국 세 건, 러시아 두 건, 예멘 한 건입니다.

앞의 네 건은 무기 자체의 소프트웨어를 만든 사례입니다. 예멘 북부의 유도무기 개발 조직은 상용 휴대폰급 비행 컴퓨터를 쓴 유도 로켓과 사거리 2,000km 이상을 목표로 한 다단 탄도미사일을 포함해 세 개 프로그램을 돌렸습니다. 중국 쪽 두 건은 어뢰 요격 체계의 화력통제 규격서 작성과 전자전·방공 제압용 표적 소프트웨어였고, 러시아 쪽 한 건은 자율 군용 드론 군집이었습니다.

뒤의 두 건은 조달과 정보 수집입니다. 러시아 방산 고객을 위해 이중용도 물자를 구했고, 지향성 에너지 무기와 그 공급사에 관한 공개 정보를 모았습니다.

그러면 이들이 AI로 실제로 얻은 것은 무엇이었을까요?

공통점은 방법 쪽에 있습니다. 이들은 이미 하드웨어에 대한 전문성과 접근권을 가진 쪽이었고, 작업을 여러 세션으로 쪼개 프로그램 전체 성격을 감췄습니다. 앤트로픽은 고위력 폭발물과 무기 개발 관련 트래픽을 걸러내는 분류기를 새로 넣었다고 밝혔습니다.

■ 이 보고서가 정말 말하는 것

사이버 챕터의 한 문장이 보고서 전체의 요약에 가깝습니다. "정교한 공격에 더 이상 정교한 공격자가 필요하지 않다"는 것입니다.

앤트로픽은 AI가 잘 자원을 갖춘 국가 지원 작전과 개인 운영자 사이의 인력·도구 격차를 무너뜨렸다고 봅니다. 훔친 API 키를 쓴 해킹 주의자 한 명, 흩어진 금전 목적 개인들, 국가 첩보 조직이 각각 여러 피해자를 상대로 캠페인을 유지했는데, 1년 전이라면 숙련된 인력 여럿과 전문 지식이 필요했을 일이라는 것입니다.

위협 조사자 입장에서 달라진 점은 여기서 나옵니다. 공격의 정교함이 배후를 추정하는 단서로서 더는 믿을 만하지 않게 됐다는 것입니다. 보고서가 능력 증가를 '속도·규모·깊이' 세 갈래로 나눠 재는 것도 이 때문입니다. 정교함으로 배후를 읽던 관행의 만료.

방어 쪽 변화도 함께 적혀 있습니다. 클로드는 이제 내부 추론을 요약해 답하기 때문에 전사를 훔쳐도 다른 모델 학습에 쓸 값이 떨어집니다. 페이블 5.1에서는 신규 API 계정이 추론 앞의 시스템 프롬프트와 도구, 메시지를 바꾸지 못하게 막았습니다. 생물학 쪽에서는 페이블 5를 이중용도 연구 질의 접근을 제한하는 더 강한 안전장치와 함께 출시했다고 밝혔습니다.

보고서에 붙은 GTG라는 표시는 'Generative Threat Groups'의 약자로, 앤트로픽이 악용 행위자에게 붙이는 내부 식별자입니다. 사례마다 계정을 차단하고 조사 결과를 안전장치에 반영했으며, 수사 당국과 다른 AI 기업에 공유했다고 설명했습니다.

Anthropic published a threat intelligence report on September 10 collecting cases in which its Claude models were misused.

It is titled "Countering misuse of AI: September 2026" and is the fourth such report, following ones in March, August and November 2025. It covers activity disrupted between December 2025 and August 2026; Korean coverage puts its length at 154 pages.

The report divides the activity into seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation. Claude Haiku, Sonnet and Opus were the models used; Fable and Mythos-class models appear in no case except one illicit distillation case.

■ Korean coverage clustered on weapons and hacking

Did Korean coverage carry all seven areas evenly?

I matched four Korean articles against the report's table of contents. Three mentioned conventional weapons more than once, and three mentioned cyber. Influence operations appeared in one, model distillation in one, and scams in none.

That would not matter if the missing parts were side notes. But influence operations is the report's longest chapter, and its nine cases are the most of any section.

So this piece goes in the order of what is missing, rather than repeating the missile and hacking details Korean outlets already covered closely.

■ Nine influence operations, aimed at elections

Anthropic defines influence operations as efforts to manipulate the information environment - political, civic and public discourse - with intent to deceive or distort, while concealing the activity's origin and sponsorship.

The nine cases originated in Russia, Iran, Turkey and across the Gulf, South Asia, Africa and Europe, and targeted audiences on six continents. Behind them were governments, state media and propaganda institutions, private firms selling influence to paying clients, domestic political operators, and in one case an opposition movement in exile.

Several were timed to elections. Russian state media produced fabricated claims about Moldova's president before the September 2025 vote, and a pro-government operator in Kenya prepared fake grassroots social posts ahead of Kenya's 2027 general election.

The most interesting passage is where Anthropic explains its own vantage point. A social media site usually sees an operation once the content is already circulating. An AI company may see it while the operation is still being built, because the planning, target selection and drafting happen inside the model.

Surveillance sits in the same chapter. China-based cases included surveillance and recruitment targeting Uyghurs in Syria, religious-affairs intelligence covering Catholic, Tibetan Buddhist and Falun Gong communities, transnational tracking under "stability maintenance," and public opinion monitoring. A mass interception platform for Mali's state intelligence service and two Iran-nexus actors building surveillance systems and a malicious Firefox extension also appear.

■ Distillation: five companies, 190 million exchanges

The chapter Korean coverage almost skipped is model distillation. Distillation itself is a legitimate training method - a larger teacher model's responses are used to train a smaller student model at lower cost.

What the report targets is illicit distillation: an industrial-scale, covert campaign to extract a model's capabilities and replicate them elsewhere without authorization, enabled by thousands of accounts built on false identities, stolen cards and stolen API keys.

The per-company figures are in the report but absent from the Korean articles. Alibaba accounted for over 151 million exchanges between May and July 2026, and Moonshot over 23 million in the same window. DeepSeek accounted for over 12.1 million across 14 days in July, Zhipu over 3.4 million across 17 days in June and July, and Xiaomi over 400,000 across 20 days in March and April. Anthropic has disrupted attacks from seven labs based in China since its first disclosure in February 2026.

Does any of this reach someone who never uses Claude?

One method matters directly to readers here. Moonshot silently forwarded customer requests to Claude instead of processing them with Kimi, then displayed Claude's responses. Users believed they were using Kimi. Over a ten-day period almost 300,000 requests were relayed, the vast majority routed to Opus.

DeepSeek did the same, and additionally tagged requests arriving through coding harnesses such as Claude Code and the Claude Agent SDK, then relayed those users' requests to Opus. DeepSeek's customers likely did not know their requests were going to Anthropic.

So the report frames this as a user-data question as much as an intellectual property one. DeepSeek, Xiaomi and Moonshot fed conversations between their own models and their users into Claude, and some of those exchanges contained sensitive information from individual users and major multinational companies.

■ Where the six weapons cases came from

This is the chapter Korean outlets covered most. The six cases break down as three in China, two in Russia and one in Yemen.

Four involved building software for the weapons themselves. A cell in northern Yemen ran three programs, including a guided rocket using a commodity phone-class flight computer and a multi-stage ballistic missile with a stated range goal above 2,000km. The two China-based cases covered a fire control specification for a torpedo interception system and targeting software for electronic warfare and air-defence suppression; the Russian case was an autonomous military drone swarm.

The other two were procurement and intelligence: sourcing dual-use goods for Russian defense customers, and collecting public information on a directed-energy weapon and its suppliers.

So what did these actors actually gain from AI?

What they share is method. These actors already had expertise in and access to the hardware, and they split their work across many sessions to conceal the full nature of their programs. Anthropic says it has launched new classifiers to detect traffic related to high-yield explosives and weapons development.

■ What the report is actually saying

One sentence in the cyber chapter comes close to summarizing the whole document: sophisticated attacks no longer require sophisticated attackers.

Anthropic argues AI has collapsed the labor and tooling gap that separated well-resourced state operations from individual operators. A hacktivist using stolen API keys, scattered financially motivated individuals, and a state espionage operator each sustained multi-victim campaigns that a year ago would have required many skilled operators and specialist knowledge.

For threat investigators, that changes something concrete: sophistication has stopped being a reliable signal of who is behind an operation. It is also why the report measures capability gain along three axes - speed, scale and depth.

Defensive changes are listed too. Claude now summarizes its internal reasoning before responding, which makes stolen transcripts less useful for training another model. Fable 5.1 introduced preserved thinking, which stops new API accounts from altering the system prompt, tools or messages preceding Claude's reasoning. On the biology side, Fable 5 shipped with stronger safeguards restricting access to dual-use biological research queries.

The GTG labels throughout the report stand for Generative Threat Groups, Anthropic's internal designators for actors observed abusing AI. In each case it banned the accounts, folded the findings into its safeguards, and shared intelligence with authorities and other AI companies.