AI 모델을 고를 때 가장 먼저 보는 것이 100만 토큰당 가격입니다. 그런데 그 숫자만으로 고르면 실제 비용이 뒤집히는 경우가 있습니다.
기업들이 이미 그 사실을 확인하고 선택을 바꾸고 있습니다.
봐야 할 것은 토큰 단가가 아니라 그 일을 끝내는 데 든 총비용입니다.
■ 지금 단가는 이렇습니다
먼저 기준선을 놓겠습니다. 100만 토큰당 입력·출력 가격입니다.
- 클로드 페이블5 — 입력 10달러 / 출력 50달러
- 클로드 오푸스5 — 입력 5달러 / 출력 25달러
- GPT-5.6 솔 — 입력 4달러 / 출력 20달러 (3개월 한시 인하 적용가)
오푸스5는 페이블5의 최첨단 성능에 근접하면서 가격은 절반이라고 앤트로픽 스스로 소개한 모델입니다. GPT-5.6 솔은 최근 인하로 오푸스5보다도 낮아졌습니다.
■ 단가가 낮은 쪽이 항상 싼 것은 아니다
여기서 함정이 하나 있습니다. 같은 일을 시켰을 때 모델마다 쓰는 토큰 양과 걸리는 시간이 다릅니다.
오픈AI는 GPT-5.6 발표문에서 솔이 코딩 에이전트 지수 평가에서 클로드 페이블5보다 출력 토큰을 절반 이하로 쓰고 시간도 절반 이하로 줄였으며 비용은 약 3분의 1 낮았다고 밝혔습니다.
이때의 비용은 API 단가가 아닙니다. 독립 평가 사이트는 코딩 에이전트 리더보드에서 비용을 작업당 평균 종량제 API 비용으로 측정한다고 설명합니다. 100만 토큰 가격이 높아도 재시도가 줄고 출력이 짧으면 실제 업무 비용은 달라집니다.
단가표는 재료값이고, 작업당 비용은 완성된 요리값입니다.
■ 실제 지출은 어디로 갔나
시장은 이미 움직였습니다. 결제 서비스 업체 램프가 미국 기업 7만 곳의 지출을 분석한 결과, 지난달 기업이 산 앤트로픽 전체 토큰 가운데 최고 성능 모델 페이블5의 비중은 약 6%, 지출액 기준으로는 11.4%였습니다.
절반 값인 오푸스5는 출시 한 달도 되지 않아 기업 지출액 기준으로 페이블5를 넘어섰습니다. 최신·최고 모델이 나오면 빠르게 주력이 되던 흐름이 바뀐 것입니다.
저렴한 쪽으로 선택지를 넓히는 움직임도 나옵니다. 램프에 따르면 지난달 AI를 쓰는 기업 가운데 오픈소스나 중국산 모델을 이용할 수 있는 '모델 서빙 플랫폼'을 쓴 기업 비중은 6.1%로 전달보다 0.2%포인트 늘었습니다.
■ 그러면 어떻게 고를 것인가
업무 성격에 따라 봐야 할 숫자가 다릅니다.
- 챗봇·개발 보조·보안 점검처럼 짧은 호출이 반복되는 일 — 토큰 단가가 총비용을 좌우합니다
- 복잡한 코딩 에이전트·장시간 자동화 — 성공률과 재시도 횟수, 출력 토큰이 총비용을 좌우합니다
- 정답이 하나인 단순 작업 — 최상위 모델을 쓸 이유가 적습니다
앤트로픽에 10억 달러를 투자한 벤처캐피털 액셀의 마일스 클레멘츠는 “대부분의 사람은 최첨단 모델을 활용할 필요가 없다”며 “고객들이 가장 발전된 모델만 선택해온 흐름은 지속 가능하지 않았다”고 말했습니다.
■ 켜기 전에 정할 것
값을 고르는 것만큼 중요한 것이 한도를 정하는 일입니다. 같은 작업을 자동으로 반복하게 두면 단가와 무관하게 비용이 쌓입니다.
미리 정해두면 좋은 것은 셋입니다.
- 언제 멈출지 — 실행 횟수나 기간의 상한
- 얼마나 쓰면 알림을 받을지 — 예산 경보 지점
- 실패하면 어디로 넘길지 — 대체 모델이나 예비 경로
단가표는 출발점이지 답이 아닙니다. 같은 일을 두 모델에 시켜 보고 끝날 때까지의 총비용을 재보는 편이 정확합니다.
The first thing most people look at when choosing an AI model is the price per million tokens. Pick on that number alone and the real cost can invert.
Companies have already noticed and are changing what they buy.
What matters is not the unit price but the total cost of finishing the job.
The current list prices
A baseline first — input and output per million tokens.
- Claude Fable 5 — $10 input / $50 output
- Claude Opus 5 — $5 input / $25 output
- GPT-5.6 Sol — $4 input / $20 output (with a three-month discount applied)
Anthropic itself presented Opus 5 as approaching Fable 5's frontier capability at half the price. GPT-5.6 Sol, after a recent cut, now sits below even Opus 5.
The cheaper unit price is not always cheaper
Here is the trap: given the same job, models differ in how many tokens they consume and how long they take.
In its GPT-5.6 announcement, OpenAI said Sol used less than half the output tokens and less than half the time of Claude Fable 5 on a coding agent index, at roughly a third of the cost.
That cost is not the API list price. The independent evaluator measures cost on its coding agent leaderboard as average pay-as-you-go API cost per task. A higher per-million price can still mean lower real cost if retries fall and output is shorter.
A price list is the cost of ingredients. Cost per task is the price of the finished dish.
Where the spending actually went
The market has already moved. Analysing spending at 70,000 US companies, the payments firm Ramp found that of all Anthropic tokens enterprises bought last month, the top model Fable 5 accounted for about 6% — and 11.4% by spend.
Opus 5, at half the price, passed Fable 5 on enterprise spend in under a month. The old pattern, where the newest and best model quickly became the workhorse, has changed.
There is movement toward cheaper options too. Ramp found 6.1% of AI-using companies used a 'model serving platform' — offering open-source or Chinese models — last month, up 0.2 percentage points.
So how should you choose
The number to watch depends on the kind of work.
- Chatbots, coding assistance, security checks — short repeated calls, where unit price drives total cost
- Complex coding agents and long-running automation — success rate, retries and output tokens drive total cost
- Simple tasks with one right answer — little reason to use the top tier
Miles Clements of Accel, which invested $1bn in Anthropic, said most people do not need to use the most advanced model, and that the pattern of customers only ever choosing the most advanced one was not sustainable.
Settle this before you switch it on
Setting a limit matters as much as picking a price. Leave the same task repeating automatically and cost accumulates regardless of the unit price.
Three things are worth deciding up front.
- When it stops — a cap on runs or duration
- When you get warned — a budget alert threshold
- Where it falls back — an alternate model or path on failure
A price list is a starting point, not an answer. Running the same job on two models and measuring the total cost to completion is more accurate.
Sources · Kukmin Ilbo · AI Times