앤트로픽과 오픈AI가 22일(현지시간) 같은 날 새 모델을 내놨습니다.

앤트로픽은 클로드 오퍼스 5.5를, 오픈AI는 GPT-6 솔과 GPT-6 루나를 공개했습니다. 국내 보도는 이를 '가성비 경쟁'으로 묶었습니다.

그런데 양사 발표문을 나란히 놓으면 두 회사가 서로를 자기 그래프에 넣고 '우리가 더 싸다'고 말하고 있습니다. 둘 다 맞습니다. 비교 대상이 서로 다르기 때문입니다.

■ 가격표를 한 자리에 모으면

먼저 공개된 단가입니다. 100만 토큰 기준입니다.

클로드 오퍼스 5.5는 입력 4달러, 출력 20달러입니다. 캐시 읽기는 0.20달러입니다. 최대 2.5배 빠른 패스트 모드는 입력 8달러, 출력 40달러로 따로 매겨져 있습니다.

오픈AI 쪽은 셋입니다. 이달 초 나온 최상위 GPT-6 아스트라가 입력 10달러, 출력 50달러이고, 이번에 나온 GPT-6 솔이 입력 2달러, 출력 10달러, GPT-6 루나가 입력 0.10달러, 출력 0.50달러입니다. 캐시 입력은 아스트라 1달러, 솔 0.20달러, 루나 0.01달러입니다.

여기서 두 가지가 눈에 띕니다. 오퍼스 5.5의 단가는 GPT-6 솔의 정확히 두 배이고, GPT-6 아스트라의 절반 이하입니다. 그러니까 두 회사의 신제품은 같은 칸에 놓인 물건이 아닙니다. 반면 캐시 읽기는 오퍼스 5.5와 솔이 0.20달러로 같습니다. 단가는 두 배, 캐시는 동률.

■ 20%와 40%는 다른 것을 잰 숫자다

오퍼스 5.5를 두고 국내 보도에 20%와 40%가 함께 나옵니다. 어느 쪽이 맞을까요?

둘 다 앤트로픽이 발표문에 쓴 숫자이고, 재는 대상이 다릅니다. 입력 4달러와 출력 20달러는 오퍼스 5보다 20% 낮은 값입니다. 이건 토큰 한 개당 가격입니다. 40%는 기본 설정에서 일반적인 작업을 돌렸을 때의 운영 비용입니다.

차이는 토큰을 얼마나 쓰느냐에서 납니다. 앤트로픽은 오퍼스 5.5가 토큰당 가격도 낮고 작업당 쓰는 토큰 수도 적어서, 합치면 비용이 40% 떨어진다고 설명했습니다. 여기에 에이전트와 코딩 작업 비용의 대부분을 차지하는 캐시 읽기가 0.20달러로 오퍼스 5보다 60% 내려간 것이 더해집니다.

속도도 같이 올렸습니다. 출력 생성이 오퍼스 5보다 30% 이상 빠르다고 밝혔습니다. 구독 요금제 쪽에서는 프로·맥스·팀과 좌석 기반 엔터프라이즈의 5시간 사용 한도를 올렸습니다.

오픈AI 쪽 50%도 같은 식으로 읽어야 합니다. 솔과 루나의 API 가격을 GPT-5.6 프로모션 가격 대비 50% 낮췄다는 것이고, 기준이 정가가 아니라 직전 프로모션가입니다.

■ 양쪽 다 상대를 자기 그래프에 넣었다

그럼 성능은 어느 쪽이 앞설까요?

흥미로운 건 벤치마크 쪽입니다. 두 발표문 모두 상대 회사 모델을 비교 대상으로 올려놨습니다.

앤트로픽은 코드 변경이 실제로 병합될 만한지 재는 프론티어코드에서 오퍼스 5.5가 기본 노력 수준으로 54.6%를 기록해 GPT-6 아스트라의 최고 점수 53.3%를 넘었고, 작업당 비용은 약 5분의 1이라고 밝혔습니다. 터미널벤치 4.0에서는 아스트라와 같은 수준을 약 40% 비용으로 냈다고 했습니다.

오픈AI는 반대편에서 같은 일을 했습니다. 업무 워크플로를 재는 오토메이션벤치에서 GPT-6 솔이 높은 노력 수준으로 클로드 오퍼스 5의 최대 노력 결과를 앞섰는데 작업당 비용은 오퍼스 5의 9%였다고 밝혔습니다. 전문 업무를 재는 에이전츠 라스트 이그잼에서는 솔이 56.4%로 오퍼스 5의 최고 점수를 넘으면서 작업당 비용은 60% 낮았다고 했습니다.

소프트웨어 공학 쪽 딥SWE에서는 솔이 68.8%로 클로드 페이블 5의 최고 점수 69.9%에 1.1%포인트 못 미쳤는데, 작업당 비용은 약 80% 낮았다고 적었습니다.

■ 정작 신모델끼리 붙인 숫자는 없다

두 주장을 겹쳐 놓으면 빈자리가 보입니다.

앤트로픽이 상대로 세운 건 GPT-6 아스트라입니다. 오픈AI의 최상위 모델이고 단가가 입력 10달러인 제품입니다. 반면 오픈AI가 상대로 세운 건 클로드 오퍼스 5와 페이블 5입니다. 둘 다 이번에 나온 오퍼스 5.5의 전작입니다.

그러니까 양쪽 다 상대의 '이번 신제품'이 아니라 다른 칸의 제품과 견줬습니다. 이번에 같은 날 나온 오퍼스 5.5와 GPT-6 솔을 같은 조건에서 붙인 숫자는 두 발표문 어디에도 없습니다. 값을 나란히 놓을 수 있는 건 단가뿐이고, 그건 솔이 절반입니다. 같은 조건의 직접 비교는 공백.

벤치마크 비교를 읽을 때 노력 수준을 함께 봐야 하는 이유이기도 합니다. 앤트로픽은 오퍼스 5.5의 기본 노력과 상대의 최고 점수를 견줬고, 오픈AI는 솔의 높은 노력이나 최대 노력을 상대의 최대 노력과 견줬습니다. 같은 줄에 놓인 두 점이 같은 조건이 아닐 수 있습니다.

■ 안전 쪽에 붙은 숫자

가격만 움직였을까요?

가격 경쟁 기사에서 빠지기 쉬운 대목이 하나 있습니다. 앤트로픽은 이번에 통제 경계를 우회하려는 성향을 재는 평가를 새로 만들었다고 밝혔습니다.

그 평가에서 오퍼스 5.5가 경계를 우회하려 시도한 빈도가 오퍼스 5나 클로드 미토스 5.1보다 약 85% 낮았다고 했습니다. 확인된 시도는 모두 심각도가 낮았고 모델이 스스로 보고했다는 설명이 붙었습니다.

오픈AI도 솔과 루나가 아스트라에서 시작한 정렬 작업을 이어받아, GPT-5.6 계열보다 오해를 부르는 진술 비율이 낮아졌다고 밝혔습니다.

두 회사가 속도 조절을 말하면서 같은 날 가격을 내린 것을 두고 국내 보도가 '안전을 외치면서 가격 경쟁은 가속'이라고 짚은 것도 이 지점입니다. 다만 발표문 기준으로는 성능·가격과 안전 평가가 함께 움직였습니다.

Anthropic and OpenAI released new models on the same day, September 22.

Anthropic announced Claude Opus 5.5; OpenAI announced GPT-6 Sol and GPT-6 Luna. Korean coverage grouped the two under a price war.

Put the two announcements side by side, though, and each company has placed the other on its own chart and claimed to be cheaper. Both are right - because they are comparing against different things.

■ The price sheets in one place

First the published rates, per million tokens.

Claude Opus 5.5 is $4 input and $20 output, with cache reads at $0.20. A fast mode running up to 2.5x speed is priced separately at $8 and $40.

OpenAI has three. GPT-6 Astra, released earlier this month as the flagship, is $10 input and $50 output. The new GPT-6 Sol is $2 and $10, and GPT-6 Luna is $0.10 and $0.50. Cached input runs $1 for Astra, $0.20 for Sol and $0.01 for Luna.

Two things stand out. Opus 5.5's rate is exactly double GPT-6 Sol's and less than half GPT-6 Astra's - so the two companies' new products are not in the same bracket. Cache reads, on the other hand, are identical at $0.20 for both Opus 5.5 and Sol.

■ 20% and 40% measure different things

Korean coverage of Opus 5.5 carries both 20% and 40%. Which is right?

Both are Anthropic's own figures, measuring different things. The $4 and $20 rates are 20% below Opus 5 - that is the per-token price. The 40% is the cost of running typical workloads at default settings.

The gap comes from token consumption. Anthropic says Opus 5.5 costs less per token and also uses fewer tokens per task, which nets out to a 40% drop. On top of that, cache reads - which make up the majority of agentic and coding costs - fell 60% to $0.20.

Speed moved too: output generation is more than 30% faster than Opus 5. On subscriptions, five-hour usage limits went up on Pro, Max, Team and seat-based Enterprise plans.

OpenAI's 50% should be read the same way. It refers to cutting API prices for Sol and Luna by half against GPT-5.6 promotional pricing - the baseline is the previous promotional rate, not list price.

■ Each put the other on its own chart

So which one actually performs better?

The benchmarks are where it gets interesting: both announcements include the rival's models.

Anthropic reports that on FrontierCode, which measures whether an agent's code changes would be merged, Opus 5.5 at default effort scored 54.6%, beating GPT-6 Astra's top score of 53.3% at roughly a fifth of the cost per task. On Terminal-Bench 4.0 it says Opus 5.5 matches Astra for about 40% of the cost.

OpenAI did the same from the other side. On AutomationBench, which tests business workflows, GPT-6 Sol at xhigh effort outperformed Claude Opus 5 at max effort at just 9% of Opus 5's cost per task. On Agents' Last Exam, Sol at max effort scored 56.4%, above Opus 5's best in that evaluation, at 60% lower cost per task.

On DeepSWE, Sol at max effort scored 68.8%, within 1.1 percentage points of Claude Fable 5's best of 69.9%, at approximately 80% lower cost per task.

■ Nobody benchmarked the two new models against each other

Overlay the two claims and a gap appears.

Anthropic's comparison target is GPT-6 Astra - OpenAI's flagship, at $10 input. OpenAI's targets are Claude Opus 5 and Fable 5, both predecessors of the Opus 5.5 announced that same day.

So each company measured itself against something other than the rival's actual new release. No figure anywhere in the two announcements puts Opus 5.5 and GPT-6 Sol on the same footing. The only numbers that line up are the rates - and there, Sol is half.

It is also why effort levels matter when reading these charts. Anthropic compared Opus 5.5 at default effort against rivals' top scores; OpenAI compared Sol at high or max effort against rivals at max effort. Two points on the same line may not be on the same settings.

■ The numbers on the safety side

Did only the prices move?

One part tends to drop out of price-war coverage. Anthropic says it built a new evaluation measuring a model's propensity to cross containment boundaries.

In it, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1. The attempts that were found were all low severity, and the model reported them itself.

OpenAI says Sol and Luna build on the alignment work introduced with Astra, showing lower rates of misleading statements than their GPT-5.6 counterparts.

Korean outlets framed two companies preaching a slower pace while cutting prices on the same day as safety talk alongside an accelerating price war. On the announcements' own terms, though, the performance, price and safety figures moved together.