오픈AI가 GPT-5.6 값을 내렸다.
가장 싼 Luna 는 80%, 중간급 Terra 는 20%. 2026년 7월 30일 발표다.
80%는 좀 센 숫자다.
오픈AI 설명으로는 Luna 가 1년 전 프런티어급 모델과 비슷한 성능을 내면서, 과제당 비용은 그 값의 6% 수준이고 속도는 9배 가까이 빠르다고 한다.
Agents' Last Exam 이라는 평가에서는 Luna 가 Fable 5 를 앞섰는데, 과제당 추정 비용은 99% 가까이 낮았다고 밝혔다.
이 값은 API 뿐 아니라 Codex 와 ChatGPT Work 에서 사용량을 구독분에 얼마나 깎아 쓰는지에도 반영된다.
그런데 같은 발표에 반대 방향 얘기가 하나 붙어 있다.
'Fast 모드'가 새로 생겼다. 기존 Priority Processing 을 대체한다.
Sol 기준으로 표준 처리보다 최대 2.5배 빠른데, 값은 두 배다. 지능은 그대로라고 못박았다.
값은 내리고, 속도는 따로 판다는 얘기다.
싸진 건 '같은 일을 더 싸게', 비싸진 건 '같은 일을 더 빨리'. 둘은 다른 상품이다.
어떻게 내릴 수 있었는지도 적어놨다.
모델·추론 시스템·에이전트 하네스를 다 손봤다는 건데, 눈에 걸린 건 이 대목이다.
사람이 관리하는 과정 안에서 GPT-5.6 Sol 이 직접 프로덕션 커널을 다시 짜고 실험을 수백 번 돌렸고, 그 결과 모델을 굴리는 비용이 20% 줄고 토큰 생성 효율이 15% 넘게 올랐다고 한다.
AI 가 자기 운영비를 깎는 데 투입됐다는 뜻으로 읽힌다.
하루 전날인 7월 29일엔 xAI 가 음성 모델을 내놨다.
Grok Voice Think Fast 2.0. 말을 듣고 바로 말로 답하는 모델이다.
여기도 결국 값과 속도 얘기였다.
- ① 가격 — 오디오 1분에 0.08달러. 예측 가능하고 투명해야 한다며 분당 단가로 못박았다
- ② 속도 — 첫 소리가 나오기까지 0.70초. 이전 버전 1.25초에서 줄었다
- ③ 방식 — 말하면서 동시에 생각한다. 그래서 더 똑똑해져도 지연이 늘지 않는다는 게 xAI 설명
- ④ 효율 — 응답당 추론 토큰이 이전 버전의 0.4배
받아쓰기 정확도는 전용 받아쓰기 모델보다 낫다고 했다. 24개 언어 짧은 문장 수천 개로 재봤더니 Deepgram Nova 3·ElevenLabs Scribe v2 대비 1.5~2배, 시끄러운 환경에서는 격차가 10배쯤으로 벌어졌다는 것.
다만 xAI 가 올린 표를 그대로 보면 전부 1등은 아니다.
대화를 주고받는 흐름을 재는 Full Duplex Bench 에서는 GPT-Realtime-2.1 이 95.7%로 Grok 의 95.1%보다 높다. 수치 출처는 Artificial Analysis 라고 표기돼 있다.
자기 발표에 자기가 진 항목을 남겨둔 건, 그래도 봐줄 만하다.
쓰는 사람 입장에서 당장 걸리는 건 날짜 하나다.
2026년 8월 5일에 grok-voice-latest 가 1.0 에서 2.0 으로 자동으로 넘어간다. 그대로 두면 알아서 바뀌고, 1.0 을 유지하려면 그 전에 grok-voice-think-fast-1.0 으로 고정해둬야 한다.
두 발표를 나란히 놓으면 공통점이 보인다.
이번엔 둘 다 '더 똑똑해졌다'가 앞이 아니었다. 값이 얼마고 몇 초 걸리는지가 앞이었다.
그래서 이게 나랑 무슨 상관이냐.
내가 쓰는 앱 상당수가 이 API 위에 올라가 있다. 원가가 내려가면 무료로 풀리는 한도나 응답 속도로 돌아오는 경우가 많다.
다만 개인 구독 요금이 같이 내린다는 얘기는 이번 발표에 없다. 값이 내렸다고 내 결제 금액이 바뀌는 건 아니라는 뜻이다.
직접 API 를 붙여 쓰고 있다면 얘기가 다르다. 같은 작업을 더 싼 모델로 내릴 수 있는지 한 번 재볼 만한 시점이다.
성능 경쟁이 끝났다는 얘기는 아닐 거다.
다만 자랑할 게 성능밖에 없던 시기는 지난 것 같다..
OpenAI cut the price of GPT-5.6.
Luna, the cheapest model, is down 80%; Terra, the mid-tier one, is down 20%. Announced July 30, 2026.
80% is a big number.
OpenAI says Luna matches models that were frontier-class a year ago, at roughly 6 cents on the dollar per task and nearly nine times the speed.
On an evaluation called Agents' Last Exam, Luna outperformed Fable 5 at an estimated cost per task nearly 99% lower.
The new prices also change how usage counts against paid subscriptions in Codex and ChatGPT Work.
But the same announcement carries something pointing the other way.
There is a new Fast mode, replacing Priority Processing.
For Sol it runs up to 2.5x faster than standard processing at twice the price. Intelligence is unchanged, OpenAI notes explicitly.
Prices go down; speed is sold separately.
What got cheaper is the same work for less. What got pricier is the same work, faster. They are two different products.
OpenAI also explains how the cut was possible.
It reworked the models, the inference systems, and the agentic harness. One line stands out.
Inside a human-led process, GPT-5.6 Sol rewrote and optimized production kernels and ran hundreds of experiments, cutting the end-to-end cost of serving the model by 20% and raising token-generation efficiency by more than 15%.
Reads to me like the AI was put to work shaving its own running costs.
A day earlier, on July 29, xAI shipped a voice model.
Grok Voice Think Fast 2.0 — it listens and answers in speech.
This one was about price and speed too.
- 1. Price — $0.08 per minute of audio. xAI says pricing should be predictable and transparent, so it set a flat per-minute rate
- 2. Speed — 0.70s to first audio, down from 1.25s in the previous version
- 3. Method — it reasons while speaking, so xAI says added intelligence costs no extra latency
- 4. Efficiency — reasoning tokens per response are 0.4x the previous version
On transcription it claims to beat dedicated speech-to-text models: across thousands of short phrases in 24 languages, a 1.5-2.0x improvement over Deepgram Nova 3 and ElevenLabs Scribe v2, widening to roughly 10x in noisy settings.
Read xAI's own table, though, and it does not win everything.
On Full Duplex Bench, which measures conversational give-and-take, GPT-Realtime-2.1 scores 95.7% against Grok's 95.1%. The numbers are credited to Artificial Analysis.
Leaving a losing row in your own announcement counts for something.
For anyone building on it, one date matters.
On August 5, 2026, grok-voice-latest moves from 1.0 to 2.0 automatically. Do nothing and it switches; to stay on 1.0, pin grok-voice-think-fast-1.0 before then.
Put the two announcements side by side and the shared thread shows.
Neither led with getting smarter. Both led with what it costs and how many seconds it takes.
So what does this mean for you?
Many of the apps you use sit on these APIs. When the underlying cost drops, it often comes back as a larger free tier or a faster response.
That said, nothing in this announcement says consumer subscription prices are coming down. A cheaper API does not mean a smaller bill for you.
If you call the API yourself, it is a different story. This is a good moment to test whether the same job runs on a cheaper model.
This is not the end of the race for raw capability.
But the stretch where capability was the only thing worth bragging about seems to be over..
Sources · OpenAI News · xAI News