앤트로픽의 클로드가 8월 24일(현지시간) 또 오류를 냈습니다.
이번이 처음이 아닙니다. 8월 16일에도 비슷한 장애가 있었고, 8일 만의 재발입니다.
반복되는 서비스 불안정에 이용자들의 우려가 커지고 있습니다.
■ 네 개 모델, 네 개 서비스
장애는 8월 24일 오전 5시 6분(UTC, 한국시간 오후 2시 6분)에 처음 감지됐습니다.
회사는 클로드 미토스5와 클로드 페이블5, 클로드 오푸스5, 클로드 오푸스4.8에서 요청 처리 시 오류가 발생하고 있다고 공지했습니다. 최상위부터 이전 세대까지 걸쳐 있습니다.
서비스도 넓게 걸렸습니다. 상태 페이지는 챗봇 클로드와 클로드 API, 개발자용 클로드 코드, 협업 도구 클로드 코워크 네 가지를 모두 부분 장애로 분류했습니다.
원인은 무엇이었을까요? 앤트로픽은 공지 몇 분 뒤 오류 원인을 특정했다고 밝히고 수정 작업에 들어갔습니다. 다만 정확히 무엇이 오류를 일으켰는지는 공개하지 않았습니다.
■ 9시간짜리 작업이 끊겼다
이용자 신고를 모으는 다운디텍터에도 제보가 급증했습니다. 노트북체크에 따르면 신고는 UTC 기준 오전 4시 55분, 한국시간 오후 1시 55분경 급격히 늘었습니다. 공식 감지 시각보다 앞섭니다.
피해는 오래 돌리는 쪽에 몰렸습니다. 레딧 이용자들은 클로드 코드 작업 흐름이 중단됐다는 불만을 쏟아냈고, 한 이용자는 9시간짜리 작업이 이번 장애로 중단됐다고 전했습니다.
짧게 묻고 답하는 사용은 재시도로 넘어가지만, 몇 시간을 돌리는 작업은 한 번의 오류로 처음부터입니다.
■ 인과관계가 확인되지 않은 이상 행동
기술적 장애 외에 다른 얘기도 나왔습니다. 일부 이용자는 클로드가 특정 캐릭터 모드나 언어적 세부 사항을 제대로 따르지 못하는 등 예기치 않은 행동을 보였다고 전했습니다.
다만 이건 이번 오류율 상승과 직접적인 인과관계가 확인된 것이 아닙니다. startuphub.ai 는 핵심 인프라 문제와 무관할 수 있다면서도, 고도화된 AI 모델이 정상 운영 상황에서도 여전히 예측하기 어려운 면모를 지닌다는 점을 보여준다고 짚었습니다.
확인된 것과 겹쳐 보이는 것을 구분해 둘 필요가 있습니다. 확인된 것은 오류율 상승과 부분 장애까지입니다.
■ 쓰는 쪽이 준비할 수 있는 것
그러면 이용자는 무엇을 할 수 있을까요? startuphub.ai 는 개발자들이 지수 백오프를 적용한 재시도 로직을 구축하고 '529 Overloaded' 같은 특정 오류 코드를 모니터링해야 한다고 조언했습니다.
서비스 가용성에 민감한 작업일수록 대체 모델이나 예비 전략을 마련해 둘 필요가 있다는 설명입니다. 기업 입장에서는 서비스 중단 가능성을 프로젝트 일정과 이용자 기대치에 미리 반영해 두면 충격을 줄일 수 있다고 덧붙였습니다.
여드레 간격으로 두 번. 한 번은 사고이고 두 번은 조건입니다.
Anthropic's Claude failed again on August 24, US time.
It was not the first time. A similar outage hit on August 16 — this one came eight days later.
Repeated instability is raising concern among users.
Four models, four services
The outage was first detected at 05:06 UTC on August 24.
The company said requests were erroring across Claude Mythos 5, Claude Fable 5, Claude Opus 5 and Claude Opus 4.8 — from the top tier down to the previous generation.
The blast radius was wide. The status page marked all four services as partial outages: the Claude chatbot, the Claude API, Claude Code for developers and the collaboration tool Claude Cowork.
What caused it? Anthropic said minutes after the notice that it had identified the cause and begun a fix. It did not disclose what the cause was.
A nine-hour job cut off
Reports surged on Downdetector, which aggregates user complaints. According to Notebookcheck, reports rose sharply around 04:55 UTC — earlier than the official detection time.
The damage fell hardest on long-running work. Reddit users complained that Claude Code workflows had been interrupted, and one said a nine-hour job was cut off by the outage.
Short question-and-answer use survives a retry. A job running for hours starts over.
Odd behaviour with no established link
Beyond the technical failure, something else came up. Some users reported unexpected behaviour, with Claude failing to follow particular character modes or linguistic details.
No causal link to the error spike has been established. startuphub.ai noted it may be unrelated to the core infrastructure problem, while arguing it shows advanced AI models remain hard to predict even in normal operation.
It is worth separating what is established from what merely overlaps in time. What is established is the error spike and the partial outage.
What users can prepare
So what can users do? startuphub.ai advised developers to build retry logic with exponential backoff and to monitor specific error codes such as '529 Overloaded'.
The more availability-sensitive the work, the more it needs a fallback model or contingency. For companies, building the possibility of downtime into project schedules and user expectations reduces the shock.
Twice in eight days. Once is an incident; twice is a condition.
Sources · Wikitree