클로드에 건강이나 생물 얘기를 물었다가 답이 갑자기 얕아진 적이 있다면, 기분 탓이 아니었다.
앤스로픽이 8월 7일 그 장치를 손봤다고 밝혔다.
답이 조용히 다른 모델로 넘어가고 있었다
클로드 페이블 5에는 안전 분류기가 붙어 있다. 생물학 관련 요청이 걸리면 그 질문을 페이블 5가 아니라 오퍼스 5로 넘겨서 답하게 하는 방식이다.
앤스로픽은 이걸 '폴백'이라고 부른다. 오퍼스 5는 생물학 능력이 페이블 5만큼 높지 않은 모델이다.
사용자 화면에는 아무 설명도 뜨지 않는다. 그냥 답이 얕아진다.
문제는 이 그물이 너무 넓었다는 것이다. 검사 수치를 읽어달라거나, 증상이 뭔지 묻거나, 공부하려고 물어본 것까지 함께 걸렸다.
이번에 분류기 규칙을 다시 써서 자체 테스트 기준 생물학 관련 폴백이 약 85% 줄었다고 한다.
생물학만이 아니라 전체 폴백도 함께 준다. Claude.ai 약 67%, 코워크 55%, 클로드 코드 17%, 클로드 플랫폼 7%다.
의료 종사자는 임상 관련 작업에서 도움을 더 받을 수 있게 된다는 설명도 붙었다.
왜 그렇게 넓게 막아뒀나
앤스로픽은 페이블 5를 내놓을 때 생물학 질문을 거의 전부 막아둔 상태로 출시했다. 일부러 그렇게 했다고 밝혔다.
이유는 이 모델이 일부 고난도 생물학 과제에서 전문가를 넘어서기 때문이다.
그리고 이 분야는 좋은 쓰임과 나쁜 쓰임을 가르는 선이 흐리다.
생백신이 그 예로 나온다. 막으려는 병원체를 실제로 배양해야 만들 수 있다. 치료법을 찾는 연구가 병을 일으키는 물질을 다뤄야 하는 경우도 있다.
앤스로픽은 이 모호함을 악용해 위험한 작업을 평범한 연구처럼 보이게 만드는 쪽이 실제로 있다고 봤다. 미국 정보기관의 2026년 연례 위협 평가도 근거로 들었다. 합성생물학과 유전체 편집을 포함한 생명공학 발전이 "새로운 생물학적 위협으로 이어질 수 있다"는 대목이다.
그래서 넓게 막고 시작한 뒤, 몇 주에 걸쳐 분류기의 규칙 묶음을 다시 썼다. 허용해야 할 쓰임을 하나하나 파내는 방식이었고, 사내외 전문가들에게 의견을 받아 반영했다고 한다.
다 막고 시작해서 하나씩 여는 순서였다. 반대로 했다면 모델 공개가 몇 주에서 몇 달 늦어졌을 거라는 게 앤스로픽의 설명이다.
그래도 안 열리는 문이 있다
다 푼 게 아니다. 이중용도로 분류하는 요청은 지금도 오퍼스 5로 넘어간다.
바이러스학, 독성학, 분자설계가 그 예로 명시됐다. 그래서 전문 생물학 연구나 신약 개발 용도로는 아직 못 쓴다고 앤스로픽이 직접 적었다.
오탐도 남는다. 위험도가 아주 낮은데도 안전 여유 구간에 걸려 분류기가 작동하는 경우는 계속 있을 거라고 했다.
수치는 자체 테스트 결과다. 85%도, 나머지 감소율도 앤스로픽이 스스로 재서 발표한 값이다. 외부 검증 수치는 공개되지 않았다.
그래도 방향은 분명하다. AI 회사가 안전장치를 조이는 발표만 하는 건 아니다. 너무 조였다고 판단하면 푸는 발표도 한다.
다만 어디를 풀고 어디를 잠글지 정하는 건 여전히 회사 쪽이다..
If you ever asked Claude about health or biology and the answer suddenly got shallower, that was not your imagination.
On August 7, Anthropic said it had adjusted the mechanism behind it.
Your question was quietly handed to another model
Claude Fable 5 runs behind a safety classifier. When a biology-related request trips it, the question is answered by Opus 5 instead of Fable 5.
Anthropic calls this a fallback. Opus 5 does not have the same level of biological capability as Fable 5.
Nothing appears on screen to explain it. The answer just gets thinner.
The problem was that the net was too wide. Asking to interpret lab results, understand symptoms, or learn biology for study got caught too.
Anthropic rewrote the classifier rules, and in its own testing biology-related fallbacks dropped by about 85%.
Total fallbacks fall as well: roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code and 7% on the Claude Platform.
Healthcare professionals should also get more support on clinical tasks.
Why block so much in the first place
Anthropic launched Fable 5 with almost all biology queries blocked. It says that was deliberate.
The reason is that the model can outperform experts on some highly complex biological tasks.
And in this field the line between beneficial and harmful use is blurry.
Live vaccines are the example given: making one requires growing the very pathogen it is meant to prevent. Some treatment research also requires handling the compounds that cause the disease.
Anthropic says sophisticated actors exploit that ambiguity to make dangerous tasks look like ordinary research. It cited the US Intelligence Community's 2026 Annual Threat Assessment, which says advances in biotechnology including synthetic biology and genomic editing "could lead to novel biological threats."
So it started wide and spent several weeks rewriting the classifier's rule set, carving out benign uses in detail and taking feedback from experts inside and outside the company.
Block everything first, then open it up piece by piece. Anthropic says the alternative would have delayed the model's general release by weeks or months.
Some doors stay shut
It is not fully open. Requests classified as dual-use still route to Opus 5.
Virology, toxicology and molecular design are named as examples. Anthropic states plainly that the model is not yet usable for professional biology research or drug development.
False positives remain too. Requests that are very low risk but fall inside the safety margin will still trip the classifier.
The figures are from internal testing. The 85% and the other reductions were measured and published by Anthropic itself. No third-party numbers were released.
Still, the direction is worth noting. AI companies do not only announce tightened safeguards. When they decide they went too far, they announce loosening too.
Though deciding what opens and what stays shut is still their call..
Sources · Anthropic Blog