클로드를 쓰다가 '모델이 바뀌었다'는 안내가 뜬 적이 있으신가요? 그 뒤로 답이 달라진 느낌이었다면 실제로 다른 모델이 답하고 있었을 수 있습니다.

앤트로픽은 이 현상을 페이블 5·5.1용과 오퍼스 5·5.5용 도움말 두 편으로 나눠 설명합니다. 두 문서를 합쳐 보면 규칙이 한눈에 들어옵니다.

요약하면 이렇습니다. 모든 요청에 자동 안전 검사가 돌고, 특정 영역에 걸리면 한 단계 낮은 모델이 같은 대화에서 다시 답합니다. 이 자동 전환은 기본으로 켜져 있고, 끌 수 있습니다.

■ 무엇에 걸리나

걸리는 영역은 모델마다 조금 다릅니다.

페이블 5와 5.1은 가장 넓게 막습니다. 익스플로잇이나 악성코드, 공격 도구를 만드는 것 같은 공격형 사이버보안, 바이러스학·독성학·약물 설계·분자 설계처럼 이중용도로 보는 생물학 질의의 상당 부분, 모델의 요약된 사고를 뽑아내려는 증류 시도, 그리고 분산 학습 인프라나 특정 칩용 커널 개발 같은 최첨단 AI 개발 과제 일부입니다. 앤트로픽은 이 차단이 일부러 넓게 잡혀 있다고 밝혔습니다.

오퍼스 5와 5.5는 훨씬 좁습니다. 대부분의 요청은 전환을 겪지 않는다고 적혀 있습니다. 고위험 공격형 사이버 요청은 오퍼스 4.8로 넘어가는데, 소스 코드의 취약점 점검처럼 안전한 코딩에는 계속 쓸 수 있습니다.

생물학은 두 모델이 갈립니다. 오퍼스 5.5는 페이블 5와 비슷한 생물학 분류기가 있어 이중용도 요청을 오퍼스 5로 넘기지만, 오퍼스 5는 생물학·화학·생명과학 질문에서 전환하지 않습니다. 최첨단 AI 개발 쪽 분류기도 오퍼스 5.5에만 있습니다.

하나는 전환 없이 바로 막힙니다. 모델의 내부 추론을 뽑아내려는 시도입니다. 추론을 글자 그대로 반복하거나 사고 과정 전체를 밖으로 적어 달라는 요청이 예로 나옵니다. 추론을 설명해 달라거나 개념을 가르쳐 달라는 건 괜찮다고 했습니다.

■ 내가 쓰지 않은 내용에도 걸린다

특정 주제를 묻지 않았는데도 바뀌었다면 이유가 있을 수 있습니다.

앤트로픽은 검사가 방금 보낸 메시지만 보는 게 아니라 모델이 읽는 모든 것을 본다고 적었습니다. 메모리와 커넥터에서 가져온 내용, 웹 검색 결과, 올린 파일까지입니다. 그래서 내가 입력하지 않은 내용 때문에도 차단될 수 있습니다.

■ 바뀐 뒤에 벌어지는 일

전환되면 클로드가 막힌 요청을 같은 대화에서 낮은 모델로 다시 돌립니다. 모델이 바뀌었다는 안내가 뜹니다.

그 뒤로 모델 선택기는 대화가 끝날 때까지 낮은 모델에 머뭅니다. 선택기에서 언제든 원래 모델로 돌아갈 수 있지만, 처음 걸린 요청이 대화에 남아 있으면 같은 안전장치에 다시 걸릴 수 있습니다. 앤트로픽은 되돌리기 전에 앞의 메시지를 고치는 게 대개 도움이 된다고 적었습니다.

비용도 나뉩니다. 페이블에서 출력이 나오기 전에 막히면 바로 오퍼스로 넘어가고 오퍼스 요금만 붙습니다. 출력 도중에 막히면 입력과 막히기 전까지 나온 부분은 페이블 요금, 나머지는 오퍼스 요금입니다.

■ 끄는 법과 끈 뒤

자동 전환이 싫다면 끌 수 있을까요?

설정의 '기능'에서 '메시지가 걸리면 모델 전환' 스위치를 끄면 됩니다. 클로드 코드에서는 Config의 MODEL & OUTPUT에 있습니다. 해당 모델을 처음 고를 때 켜진 상태로 시작합니다.

끄면 막힌 요청은 모델을 바꾸는 대신 대화가 멈춥니다. 그다음은 메시지를 고쳐 같은 모델로 다시 보내거나, 같은 메시지를 낮은 모델로 직접 보내는 것 중에 고릅니다.

이 규칙은 웹과 모바일, 데스크톱, 코워크, 클로드 코드에 똑같이 적용됩니다. API는 다릅니다. 자동 전환이 기본으로 꺼져 있고, 개발자가 직접 설정해야 합니다.

정당한 보안·생명과학 일이 계속 막힌다면 길이 있습니다. 방어 목적의 보안 작업은 '사이버 검증 프로그램', 생명과학 연구 조직은 '생명과학 검증 프로그램'에 신청할 수 있고, 관련 없는 주제인데 막혔다면 '피드백 보내기'로 알리라고 안내합니다.

■ 정리하면

  • '모델이 바뀌었다' 안내는 안전 검사에 걸렸다는 뜻입니다 — 한 단계 낮은 모델이 대신 답합니다
  • 오퍼스 5·5.5는 대부분 전환이 없습니다 — 가장 넓게 막는 건 페이블 5·5.1입니다
  • 내가 쓰지 않은 내용도 원인일 수 있습니다 — 메모리·커넥터·검색 결과·파일까지 검사합니다
  • 되돌리기 전에 앞 메시지를 고칩니다 — 그대로 두면 다시 걸릴 수 있습니다
  • 설정 → 기능에서 끌 수 있습니다 — 끄면 바뀌는 대신 멈춥니다
  • 추론을 그대로 뽑아 달라는 요청은 전환 없이 바로 막힙니다

Have you seen a notice in Claude saying the model switched? If answers felt different afterwards, a different model may indeed have been answering.

Anthropic explains this in two help pages, one for Fable 5 and 5.1 and one for Opus 5 and 5.5. Put together, the rules become clear.

In short: automated safety checks run on every request, and when a request trips a specific area, a less capable model re-answers it in the same conversation. This automatic switching is on by default, and you can turn it off.

■ What trips it

The areas differ slightly by model.

Fable 5 and 5.1 block most broadly: offensive cybersecurity such as building exploits, malware or attack tooling; a large fraction of dual-use biology such as virology, toxicology, drug design and molecular design; distillation attempts including extracting the model's summarized thinking; and a narrow set of frontier LLM development tasks such as distributed training infrastructure and kernel development for certain chips. Anthropic says these safeguards are intentionally broad.

Opus 5 and 5.5 are much narrower - most requests never meet a fallback. Higher-risk offensive cyber requests fall back to Opus 4.8, while secure coding such as scanning source code for vulnerabilities remains available.

Biology splits the two. Opus 5.5 has Fable-like biology classifiers that send dual-use requests to Opus 5, but Opus 5 itself does not fall back on biology, chemistry or life-science questions. The frontier LLM development classifiers apply only to Opus 5.5.

One category is blocked outright with no fallback: attempts to extract the model's internal reasoning, such as asking it to repeat its reasoning verbatim or write its full chain of thought to an external output. Asking it to explain its reasoning or teach a concept is fine.

■ Content you did not type can trigger it

If it switched without you asking about any of those topics, there may be a reason.

Anthropic says the checks review everything the model reads, not only your latest message - memory, content from connectors, web search results and files. So a block can be triggered by content you did not type.

■ What happens after the switch

Claude re-runs the blocked request on the less capable model in the same conversation, and a notice explains that the model switched.

The model picker then stays on that model for the rest of the conversation. You can switch back anytime, but if the original request is still in the conversation, the same safeguards may trigger again. Anthropic says editing your previous message before retrying often helps.

Billing splits too. If Fable blocks a request before producing any output, it switches to Opus at once and you pay Opus rates only. If it is blocked midstream, the input and tokens streamed before the block are charged at Fable rates and the rest at Opus rates.

■ Turning it off, and what follows

Can you turn automatic switching off?

Yes: in Settings > Capabilities, toggle "Switch models when a message is flagged" off - in Claude Code, under Config > MODEL & OUTPUT. It starts on the first time you select the model.

With it off, a blocked request pauses the conversation instead of switching. You then either edit your message and retry on the same model, or send the same message to a less capable model manually.

These rules apply the same way on the web, mobile, desktop, Cowork and Claude Code. The API is different: automatic switching is off by default and developers must configure fallbacks themselves.

If legitimate security or life-science work keeps getting blocked, there are routes: defensive security work can apply to the Cyber Verification Program, life-science research organizations to the Life Sciences Verification Program, and unrelated blocks can be reported with "Send feedback."

■ In short

  • A 'model switched' notice means a safety check fired - a less capable model is answering instead
  • Opus 5 and 5.5 rarely switch - Fable 5 and 5.1 block most broadly
  • Content you did not type can cause it - memory, connectors, search results and files are all checked
  • Edit the earlier message before switching back - otherwise it may trip again
  • Turn it off in Settings > Capabilities - it then pauses instead of switching
  • Requests to extract reasoning verbatim are blocked outright, with no fallback