오픈AI가 새 서비스 등급을 내놨습니다. 이름은 울트라패스트(Ultrafast)입니다.

같은 모델을 더 빠르게 돌리는 등급입니다. GPT-5.6 Sol을 표준 처리보다 최대 14배 빠르게 돌린다고 했고, API에 먼저 들어갑니다.

같은 기간 값도 함께 움직였습니다. 이번 발표는 속도 한 축만 본 것이 아닙니다.

■ 숫자는 초당 750토큰

출력 토큰이 초당 최대 750개입니다. 칩은 세레브라스(Cerebras)가 댑니다.

지금까지는 실시간 속도를 원하면 더 작거나 특화된 모델을 골라야 했습니다. 속도를 사려면 지능을 팔았다는 뜻입니다. 그 거래를 하지 않아도 된다면 무엇이 달라질까요?

모델을 바꾸는 것이 아니라는 점이 핵심입니다. GPT-5.6 Sol 그대로이고, 돌리는 속도만 다릅니다.

■ 기다리면 늦는 일들

오픈AI가 든 활용 사례는 다섯입니다. 장애 대응에서는 시스템이 죽었을 때 로그와 최근 코드 변경, 엔지니어 보고를 분석해 장애가 진행되는 동안 원인과 수정안을 뽑습니다. 금융 리서치와 보안에서는 시장 신호를 읽고 거래를 평가해 상황이 바뀌는 중에 의심 활동을 잡습니다.

고객 지원과 음성에서는 여러 단계를 거쳐야 답이 나오는 문의도 통화를 끊지 않고 실시간으로 처리합니다. 커머스에서는 상품 문의와 재고 확인, 추천, 결제 오류를 쇼핑객이 아직 고민하는 동안 해결합니다. 라이브 리서치에서는 밤새 돌려놓고 아침에 확인하던 실험을 앉은 자리에서 고쳐가며 다시 돌립니다.

공통점이 보입니다. 전부 답이 늦으면 소용없어지는 일입니다. 장애는 끝난 뒤에 원인을 알아봐야 늦고, 장바구니는 고민하는 동안 답이 없으면 닫힙니다.

■ 먼저 써본 곳들이 한 말

오픈AI는 코딩과 커머스, 금융 리서치, 고객 지원 등에서 일부 기업과 먼저 시험했다고 밝혔습니다.

“세레브라스가 가져온 속도 향상은 인상적이다. 모델을 쓰는 방식 자체를 바꾸고, 개발자가 모델과 나란히 더 집중해서 생산적으로 일할 수 있게 해준다.” — 존 크레페치, 제인스트리트

음성 쪽 얘기가 구체적입니다. 포디움의 코틀랜드 라이킨스 음성AI 프로덕트 리드는 “울트라패스트가 우리 음성 스택에서 매우 유용했다. 복잡한 작업일수록 속도가 통화 경험을 완전히 바꿔놓는다”고 했습니다.

베이시스 공동창업자 미치 트로야노프스키는 진짜 병목이 어디였는지를 짚었습니다. “정말 빠른 제품을 가로막는 장벽은 초당 토큰 수만이 아니라 모델의 지능이기도 하다. 울트라패스트는 둘을 결합한다”는 것입니다.

오픈AI 내부 개발자들도 쓰고 있다고 했습니다. 알림이 뜨면 로그를 읽고 추적을 분석해 다음 점검 항목을 뽑는 장애 대응, 그리고 연결된 도구들을 가로질러 자료를 모으고 정리하는 리서치입니다.

■ 같은 기간 값도 내려갔다

그런데 값은 어떻게 됐을까요? 속도만 움직인 것이 아닙니다. 비용 경쟁의 기준선 자체가 옮겨가고 있습니다.

오픈AI는 7월 9일 GPT-5.6 발표문에서, GPT-5.6 Sol이 Artificial Analysis 코딩 에이전트 지수에서 클로드 페이블 5보다 출력 토큰을 절반 이하로 쓰고 시간도 절반 이하로 줄였으며 비용은 약 3분의 1 낮았다고 밝혔습니다.

여기서 말한 비용은 API 단가가 아닙니다. Artificial Analysis 는 코딩 에이전트 리더보드에서 비용을 작업당 평균 종량제 API 비용으로 잰다고 설명합니다. 100만 토큰 가격이 높아도 재시도와 출력량이 줄면 실제 업무 비용은 달라진다는 뜻입니다.

단가 쪽도 내려갔습니다. 오픈AI는 7월 30일 GPT-5.6 Luna 가격을 80%, GPT-5.6 Terra 가격을 20% 낮췄습니다. 로이터는 이 인하를 기업 고객의 AI 비용 부담과 중국 저가 모델 경쟁에 대한 대응으로 전했습니다.

비교 대상은 이렇습니다. 앤트로픽은 클로드 페이블 5와 미토스 5 가격을 입력 100만 토큰당 10달러, 출력 50달러로 공지했습니다. 딥시크는 deepseek-chat 기준 입력 0.27달러에 출력 1.10달러이고, Z.AI 의 GLM-5.1은 1.4달러와 4.4달러, 문샷의 키미 K2.6은 0.95달러와 4.00달러입니다.

■ 단가표가 답을 주지 않는 이유

그러면 싼 모델을 고르면 되는 걸까요? 실제 사용 데이터는 그렇게 단순하지 않습니다.

Vercel 은 2026년 7월 AI 게이트웨이 생산 지표에서, 6월 기준 앤트로픽이 전체 지출의 61%를 차지했지만 토큰 비중은 32%였다고 밝혔습니다. 반면 오픈웨이트 모델은 토큰의 29%를 처리하면서 지출 비중은 4% 미만이었습니다. 같은 보고서에서 딥시크는 게이트웨이 토큰의 22.6%를 차지했습니다.

Artificial Analysis 의 다른 비교에서는 키미 K3가 인텔리전스 지수 60으로 GPT-5.6 Terra 보다 높았지만, GPT-5.6 Terra(high)가 100만 토큰당 1.74달러로 키미 K3의 2.31달러보다 낮고 속도도 더 빠른 것으로 제시됐습니다.

챗봇과 개발 보조처럼 반복 호출이 많은 업무는 단가가 중요하고, 복잡한 코딩 에이전트나 장시간 자동화는 성공률·재시도·출력 토큰이 총비용을 좌우합니다.

■ 지금 신청해도 못 쓴다

울트라패스트는 제한된 프리뷰입니다. 선별된 일부 고객만 씁니다. 오픈AI는 용량이 늘면 확대하겠다고 했고, 지금 열려 있는 것은 접근이 확대될 때 알려달라는 신청뿐입니다.

그래서 이번 발표는 '지금 쓸 수 있는 것'이 아니라 '방향'을 본 셈입니다. 같은 모델을 속도 등급으로 나눠 파는 방식이 자리를 잡으면, 앞으로 모델을 고르는 기준에 속도 등급이 하나 더 붙습니다.

덧붙이면, 오픈AI는 이걸 세레브라스와의 협력에서 “다음 단계”라고 표현했습니다.

속도 등급과 작업당 비용. 모델을 고를 때 보던 두 축이 같은 달에 함께 움직였습니다.

OpenAI has introduced a new service tier called Ultrafast.

It runs the same model faster — GPT-5.6 Sol at up to 14 times standard processing speed — and lands in the API first.

Price moved over the same stretch. This announcement is not only about speed.

The number is 750 tokens per second

Output runs at up to 750 tokens per second. The silicon comes from Cerebras.

Until now, wanting real-time speed meant picking a smaller or specialised model. Buying speed meant selling intelligence. What changes if you no longer have to make that trade?

The point is that the model does not change. It is GPT-5.6 Sol as it is; only the speed of execution differs.

Work where waiting defeats the purpose

OpenAI listed five uses. In incident response it analyses logs, recent code changes and engineer reports to produce causes and fixes while an outage is still running. In financial research and security it reads market signals and evaluates trades to catch suspicious activity as conditions change.

In customer support and voice it handles multi-step queries in real time without dropping the call. In commerce it resolves product questions, stock checks, recommendations and payment errors while the shopper is still deciding. In live research it reruns, with edits, experiments that used to be left overnight.

They have something in common. All of them stop being useful if the answer is late. Finding the cause after an outage ends is too late, and a cart closes while the shopper waits.

What early users said

OpenAI says it tested with selected companies in coding, commerce, financial research and customer support.

“The speedup Cerebras brings is impressive. It changes how you use the model and lets developers work alongside it more focused and more productively.” — Jon Krepechi, Jane Street

The voice case is concrete. Cortland Rykins, voice AI product lead at Podium, said Ultrafast “has been very useful in our voice stack” and that “the more complex the task, the more speed transforms the call experience.”

Mitch Troyanovsky, co-founder of Basis, named the real bottleneck: “The barrier to a truly fast product isn't only tokens per second — it's also the model's intelligence. Ultrafast combines the two.”

OpenAI's own developers use it as well: incident response that reads logs and traces on an alert to produce the next checks, and research that gathers and organises material across connected tools.

Price came down over the same stretch

And what happened to price? Speed was not the only thing moving. The baseline of the cost competition is shifting.

In its July 9 GPT-5.6 announcement, OpenAI said GPT-5.6 Sol used less than half the output tokens and less than half the time of Claude Fable 5 on the Artificial Analysis coding agent index, at roughly a third of the cost.

That cost is not the API list price. Artificial Analysis measures cost on its coding agent leaderboard as the average pay-as-you-go API cost per task. A higher per-million-token price can still mean lower real cost if retries and output volume fall.

List prices fell too. On July 30 OpenAI cut GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. Reuters framed the cuts as a response to enterprise cost pressure and competition from cheap Chinese models.

For comparison, Anthropic lists Claude Fable 5 and Mythos 5 at $10 per million input tokens and $50 per million output. DeepSeek lists deepseek-chat at $0.27 and $1.10, Z.AI's GLM-5.1 at $1.4 and $4.4, and Moonshot's Kimi K2.6 at $0.95 and $4.00.

Why the price list does not settle it

So should you just pick the cheap model? Real usage data is not that simple.

In its July 2026 AI Gateway production metrics, Vercel reported that as of June, Anthropic accounted for 61% of spend but 32% of tokens, while open-weight models handled 29% of tokens for under 4% of spend. In the same report, DeepSeek accounted for 22.6% of gateway tokens.

In another Artificial Analysis comparison, Kimi K3 scored higher on the intelligence index at 60, but GPT-5.6 Terra (high) came in at $1.74 per million tokens against Kimi K3's $2.31, and was faster.

For repetitive, high-call work like chatbots and coding assistance, unit price matters; for complex coding agents and long-running automation, success rate, retries and output tokens drive the total.

You cannot use it by signing up today

Ultrafast is a limited preview for selected customers. OpenAI says it will widen access as capacity grows; what is open now is a form asking to be told when that happens.

So this announcement is a direction rather than something you can use. If selling one model in speed tiers takes hold, choosing a model gains one more axis.

OpenAI described it as the “next step” in its work with Cerebras.

Speed tiers and cost per task. Two of the axes people use to pick a model moved in the same month.

Sources · OpenAI News · TokenPost