두 달 사이 AI 모델 값이 무너졌습니다.
7월 21일 중국 문샷AI 의 키미 K3 로 시작해 8월 23일 기업 지출 데이터로 이어지는 흐름입니다. 그 사이에 미국 선두 두 곳이 스스로 값을 내렸습니다.
마지막에 나온 숫자가 이 흐름의 성격을 바꿉니다. 기업은 이미 최고 성능 모델을 사지 않고 있었습니다.
■ 시작은 3분의 1 값이었다
문샷AI 가 7월 21일 키미 K3 를 공개했습니다. 아티피셜 애널리시스 종합 지능 지수에서 페이블5와 GPT-5.6 솔에 이어 3위였습니다.
달랐던 것은 값입니다. 과제 하나를 처리하는 평균 비용이 0.95달러로, 클로드 페이블5의 2.75달러의 3분의 1 수준이었습니다.
시장분석 기관 IG 추정으로 오픈AI와 앤트로픽의 장외 몸값이 3140억달러, 약 464조원 줄었습니다. 앤트로픽 7.31%, 오픈AI 5.62% 하락입니다. 실제로 오간 돈이 아니라 장외에서 매기는 예상 몸값이라는 점은 감안해야 합니다.
■ 사흘 뒤, 선두가 스스로 반값을 불렀다
앤트로픽이 7월 24일 클로드 오퍼스5 를 내놨습니다. 종합 지능 지수 61점으로 자사 최상위 페이블5보다 1점 높았습니다.
값은 절반이었습니다. 100만 토큰당 5~25달러로, 페이블5의 10~50달러의 정확히 반입니다.
다만 전부 이긴 것은 아니었습니다. 에이전트 코딩에서는 GPT-5.6 솔이 72.7%로 오퍼스5의 68.8%를 앞섰고, 전문가 추론은 페이블5가 56.5%로 오퍼스5의 56.3%보다 근소하게 높았습니다.
■ 8월, 값은 계속 내려갔다
오픈AI 가 7월 30일 GPT-5.6 루나를 80%, 테라를 20% 내렸습니다. 같은 발표에 반대 방향도 붙어 있었습니다. 표준 처리보다 최대 2.5배 빠른 Fast 모드를 값 두 배로 따로 팔기 시작한 것입니다.
값은 내리고 속도는 따로 판다는 뜻입니다. 같은 일을 더 싸게, 또는 같은 일을 더 빨리. 둘은 다른 상품이 됐습니다.
딥시크는 7월 31일 V4-플래시를 정식 출시했습니다. 입력 100만 토큰당 0.14달러, 출력 0.28달러입니다. 테스트 1회 평균 3센트로, 클로드 페이블5의 3.15달러와 비교하면 105분의 1이었습니다.
성능은 얼마나 떨어졌을까요? 종합 지능 지수 50점으로 구글 제미나이 3.6 플래시와 같았습니다. 상위권과는 9점 이상 벌어집니다. 값은 105배, 점수는 9점.
8월 22일에는 오픈AI 가 최상위 솔까지 3개월 한시로 20% 이상 내렸습니다. 100만 토큰당 입력 4달러, 출력 20달러가 되면서 오퍼스5보다 싸졌습니다.
■ 싼 값에 붙은 조건
모든 인하가 같은 성격이었을까요? 그렇지는 않았습니다.
메타가 8월 5일 코딩 에이전트 뮤즈 코드를 공개했습니다. 자사 벤치마크에서도 클로드 코드에 뒤진 3등이었지만, 알렉산더 왕 메타 AI 총괄은 컨트리뷰터 등급이 종량제보다 10배 넘게 저렴하다고 했습니다.
조건이 있었습니다. 그 등급에 가입하면 업계 표준 방식에 따라 모델 개선에 기여하게 된다는 설명입니다. 쓴 코드가 학습에 들어간다는 뜻입니다. 메타는 이와 별도로 제로 데이터 보유 옵션을 함께 내놨습니다.
별도 옵션을 만들었다는 것은 기본값이 그쪽이 아니라는 뜻으로 읽힙니다. 값이 싼 이유를 알고 쓰는 것과 모르고 쓰는 것은 다릅니다.
■ 수요는 이미 옮겨가 있었다
그러면 이 인하 경쟁은 공급자끼리의 싸움이었을까요? 8월 23일 나온 데이터는 다른 그림을 보여줍니다.
파이낸셜타임스가 결제 서비스 업체 램프의 데이터를 인용해, 기업이 앤트로픽 모델에 쓴 돈 가운데 페이블5 비중이 11.4%, 토큰 기준으로는 약 6%에 그쳤다고 보도했습니다. 미국 기업 7만 곳의 지출을 분석한 수치입니다.
절반 값인 오퍼스5 는 출시 한 달도 되지 않아 지출액 기준으로 페이블5를 넘어섰습니다. 최신 최고 모델이 나오면 곧 주력이 되던 이전과는 다른 양상입니다.
앤트로픽에 10억 달러를 투자한 액셀의 마일스 클레멘츠는 “고객들이 가장 발전된 모델만 선택해온 흐름은 지속 가능하지 않았다”며 “대부분의 사람은 최첨단 모델을 활용할 필요가 없다”고 말했습니다.
투자한 쪽에서 나온 말입니다. 다만 이 수치는 램프 한 곳의 결제 데이터입니다. 7만 곳이 작은 표본은 아니지만 시장 전체의 지출 구조와 같지는 않습니다.
■ 두 달이 바꾼 질문
개인 구독료가 내렸다는 발표는 이 두 달 사이에 없었습니다. API 값이 내려가도 챗GPT 플러스나 프로의 월 결제 금액은 그대로입니다.
달라진 것은 질문입니다. 두 달 전에는 어느 모델이 가장 똑똑한지를 물었습니다. 지금은 내 일에 몇 점이 필요한지를 묻게 됐습니다.
105배의 값 차이와 9점의 점수 차이. 번역이나 분류처럼 정답 판정이 쉬운 일이라면, 답은 그 사이 어딘가에 있습니다.
출처 · 매일경제 — 딥시크는 예고편…성능 비슷한데 요금 반토막 낸 ‘키미 K3’ 쇼크 · 디지털투데이 — "中 AI가 세상 지배하면…" 미국이 키미 K3 경계하는 이유 · Anthropic 공식블로그 — Introducing Claude Opus 5 · 디지털타임스 — 앤트로픽, 가성비 앞세운 AI '오퍼스5' 출시… "GPT-5.6 능가, 비용은 절반" · OpenAI 공식뉴스 — Advancing the price-performance frontier with GPT-5.6 · 데일리안 — “딥시크 새 모델, 클로드 105분의 1 비용…제미나이급 성능” · 연합뉴스 — 메타, 첫 AI 코딩 에이전트 공개…"성능 아닌 가격으로 경쟁" · AI타임스 — 오픈AI, 'GPT-5.6 솔'까지 가격 일시 인하...코덱스 사용자 2천만 돌파 · 국민일보 — 성능보다 가성비?…앤트로픽 '클로드 페이블5' 주춤
AI model prices collapsed over two months.
The run starts with Moonshot AI's Kimi K3 on 21 July and ends with corporate spending data on 23 August. In between, the two American leaders cut their own prices.
The last number changes what the run means. Enterprises had already stopped buying the top-performing models.
It began at a third of the price
Moonshot AI released Kimi K3 on 21 July. On Artificial Analysis's composite intelligence index it placed third, behind Fable 5 and GPT-5.6 Sol.
Price was the difference. The average cost of completing one task was $0.95, about a third of Claude Fable 5's $2.75.
By IG's estimate, the pre-IPO valuations of OpenAI and Anthropic fell by $314bn, roughly 464 trillion won — Anthropic down 7.31%, OpenAI down 5.62%. These are secondary-market estimates rather than money that changed hands.
Three days later, the leader halved its own price
Anthropic released Claude Opus 5 on 24 July. It scored 61 on the composite intelligence index, one point above the company's own flagship Fable 5.
The price was half: $5-25 per million tokens against Fable 5's $10-50.
It did not win everything. On agentic coding GPT-5.6 Sol led at 72.7% against Opus 5's 68.8%, and on expert reasoning Fable 5 edged ahead at 56.5% to 56.3%.
Through August, prices kept falling
On 30 July OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. The same announcement moved the other way too: a new Fast mode, up to 2.5x quicker than standard processing, sold separately at double the price.
Cheaper on one axis, sold separately on the other. The same job done cheaper, or the same job done faster. They became different products.
DeepSeek shipped V4-Flash on 31 July at $0.14 per million input tokens and $0.28 output. At three cents per test run against Claude Fable 5's $3.15, that is one hundred and fifth of the cost.
How far did performance fall? It scored 50 on the composite index, level with Google's Gemini 3.6 Flash, and nine or more points below the leaders. A 105-fold gap in price, a nine-point gap in score.
On 22 August OpenAI cut even its top Sol model by more than 20% for three months, to $4 input and $20 output per million tokens — undercutting Opus 5.
The conditions attached to cheap
Were all the cuts the same kind of thing? They were not.
Meta launched its coding agent Muse Code on 5 August. It placed third even on Meta's own benchmarks, behind Claude Code, but Alexandr Wang, Meta's head of AI, said the contributor tier costs more than ten times less than pay-as-you-go.
There was a condition. Joining that tier means contributing to model improvement in the industry-standard way, as Meta put it — your code goes into training. Meta separately introduced a zero data retention option.
Creating a separate option suggests that is not the default. Knowing why something is cheap is different from not knowing.
Demand had already moved
So was this a fight among suppliers? Data published on 23 August suggests otherwise.
The Financial Times, citing the payments company Ramp, reported that Fable 5 accounted for 11.4% of what enterprises spent on Anthropic models and about 6% of tokens. The figures come from an analysis of 70,000 US companies.
Opus 5, at half the price, passed Fable 5 on spending within a month of launch. That is not how it used to go, when each new top model quickly became the default.
Miles Clements of Accel, which has invested a billion dollars in Anthropic, said the pattern of customers always choosing the most advanced model was not sustainable, and that most people have no need for a frontier model.
That came from an investor. Still, these are one payment company's figures. Seventy thousand firms is not a small sample, but it is not the whole market either.
The question these two months changed
Nothing in these two months said consumer subscriptions were getting cheaper. Falling API prices do not change the monthly bill for ChatGPT Plus or Pro.
What changed is the question. Two months ago it was which model is smartest. Now it is how many points the job actually needs.
A 105-fold gap in price against a nine-point gap in score. For work with a clear right answer, like translation or classification, the answer sits somewhere in between.
Sources · Maeil Business Newspaper · Digital Today · Anthropic Blog · Digital Times · OpenAI News · Dailian · Yonhap News · AI Times · Kukmin Ilbo