오픈AI가 새 서비스 등급을 내놨다. 이름은 울트라패스트(Ultrafast)다.

같은 모델을 더 빠르게 돌리는 등급이다. GPT-5.6 Sol을 표준 처리보다 최대 14배 빠르게 돌린다고 했다. API에 먼저 들어간다.

숫자는 초당 750토큰

출력 토큰이 초당 최대 750개다. 칩은 세레브라스(Cerebras)가 댄다.

지금까지는 실시간 속도를 원하면 더 작거나 특화된 모델을 골라야 했다. 속도를 사려면 지능을 팔았다는 뜻이다.

속도를 얻으려고 지능을 포기하지 않아도 되면, AI가 업무에서 가장 시간에 쫓기는 자리로 들어갈 수 있다 — 오픈AI가 이번 발표에서 내세운 논리다.

모델을 바꾸는 게 아니라는 점이 핵심이다. GPT-5.6 Sol 그대로다. 돌리는 속도만 다르다.

기다리면 늦는 일들

오픈AI가 든 활용 사례는 다섯 가지다.

  • 장애 대응 — 시스템이 죽었을 때 로그·최근 코드 변경·엔지니어 보고를 분석해, 장애가 진행되는 동안 원인과 수정안을 뽑는다
  • 금융 리서치·보안 — 시장 신호를 읽고 거래를 평가해, 상황이 바뀌는 중에 의심 활동을 잡는다
  • 고객 지원·음성 — 여러 단계를 거쳐야 답이 나오는 문의도 통화를 끊지 않고 실시간으로 처리한다
  • 커머스 — 상품 문의·재고 확인·추천·결제 오류를, 쇼핑객이 아직 고민하는 동안 해결한다
  • 라이브 리서치 — 밤새 돌려놓고 아침에 확인하던 실험을, 앉은 자리에서 고쳐가며 다시 돌린다

공통점이 보인다. 전부 답이 늦으면 소용없어지는 일이다. 장애는 끝난 뒤에 원인을 알아봐야 늦고, 장바구니는 고민하는 동안 답이 없으면 닫힌다.

먼저 써본 곳들이 한 말

오픈AI는 코딩·커머스·금융 리서치·고객 지원 등에서 일부 기업과 먼저 시험했다고 밝혔다. 공개된 코멘트 중 셋.

"세레브라스가 가져온 속도 향상은 인상적이다. 모델을 쓰는 방식 자체를 바꾸고, 개발자가 모델과 나란히 더 집중해서 생산적으로 일할 수 있게 해준다." — 존 크레페치, 제인스트리트

음성 쪽 얘기가 구체적이다. 포디움의 코틀랜드 라이킨스 음성AI 프로덕트 리드는 "울트라패스트가 우리 음성 스택에서 매우 유용했다. 복잡한 작업일수록 속도가 통화 경험을 완전히 바꿔놓는다"고 했다.

베이시스 공동창업자 미치 트로야노프스키는 진짜 병목이 어디였는지를 짚었다.

"정말 빠른 제품을 가로막는 장벽은 초당 토큰 수만이 아니라 모델의 지능이기도 하다. 울트라패스트는 둘을 결합한다."

오픈AI 내부 개발자들도 쓰고 있다고 했다. 알림이 뜨면 로그를 읽고 추적을 분석해 다음 점검 항목을 뽑는 장애 대응, 그리고 연결된 도구들을 가로질러 자료를 모으고 정리하는 리서치다.

지금 신청해도 못 쓴다

제한된 프리뷰다. 선별된 일부 고객만 쓴다. 오픈AI는 용량이 늘면 확대하겠다고 했고, 지금 열려 있는 건 접근이 확대될 때 알려달라는 신청뿐이다.

그래서 이번 발표는 '지금 쓸 수 있는 것'이 아니라 '방향'을 본 셈이다. 같은 모델을 속도 등급으로 나눠 파는 방식이 자리를 잡으면, 앞으로 모델을 고르는 기준에 속도 등급이 하나 더 붙는다.

덧붙이면, 오픈AI는 이걸 세레브라스와의 협력에서 "다음 단계"라고 표현했다.

OpenAI has introduced a new service tier called Ultrafast.

It runs the same model faster — GPT-5.6 Sol at up to 14x the speed of Standard processing, the company says. It lands in the API first.

The number is 750 tokens per second

Up to 750 output tokens per second. The silicon comes from Cerebras.

Until now, getting real-time speed usually meant picking a smaller or more specialized model. Buying speed meant selling intelligence.

When speed no longer requires giving up intelligence, AI can move into the most time-sensitive parts of a business — that is the argument OpenAI makes in this announcement.

The key point is that the model does not change. It is still GPT-5.6 Sol. Only the speed it runs at is different.

Work that is useless when it arrives late

OpenAI listed five scenarios.

  • Incident response — read logs, recent code changes and engineer reports to identify the likely cause and prepare a fix while the outage is still unfolding
  • Financial research and security — analyze market signals and assess transactions to spot suspicious activity while conditions are still changing
  • Customer support and voice — resolve issues that need multiple steps or systems without interrupting the call
  • Commerce — handle product questions, inventory checks, recommendations and checkout errors while the shopper is still deciding
  • Live research — turn an experiment you used to launch overnight into something you adjust and rerun in one sitting

The pattern is clear. All of it is work that stops mattering if the answer is late. Finding the cause after the outage ends is too late, and a cart closes while the shopper waits.

What the early customers said

OpenAI says it tested with an initial group of companies across coding, commerce, financial research and support. Three of the published comments.

"The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them." — John Crepezzi, Jane Street

The voice case is the most concrete. Courtland Lykins, product lead for voice AI at Podium: "For us the Ultrafast has been invaluable in our voice stack. The speed completely changes the call experience for the more complex work."

Mitch Troyanovsky, co-founder of Basis, named where the real bottleneck was.

"Oftentimes the barrier to truly fast products is not just tokens per second, but also model intelligence, and ultrafast combines both."

OpenAI's own developers are using it too — incident response, where an alert means reading logs, analyzing traces and identifying the next checks, and research, where it gathers and organizes information across connected tools.

Signing up will not get you in

It is a limited preview for a select group of customers. OpenAI says access will expand as capacity grows, and what is open right now is only a form to be notified when it does.

So this is a look at a direction rather than something to use today. If selling one model in speed tiers sticks, picking a model gains one more axis to choose along.

OpenAI also frames this as the next step in its partnership with Cerebras.

Sources · OpenAI News