딥시크의 새 모델이 클로드의 105분의 1 비용으로 돌아간다고 한다.

AI 성능 분석업체 아티피셜 애널리시스가 현지시간 3일 내놓은 벤치마크 결과를 로이터가 전했고, 데일리안이 이를 인용했다.

모델은 'V4-플래시'. 딥시크가 7월 31일 정식 출시한 API 버전이다.

값부터 보면 이렇다.

  • 입력 100만 토큰당 0.14달러 · 출력 100만 토큰당 0.28달러
  • 테스트 1회당 평균 3센트
  • 비교 — 키미 K3 86센트 · GPT-5.6 솔 1.86달러 · 클로드 페이블 5 3.15달러

3센트와 3.15달러. 105배 차이가 여기서 나온다.

그럼 성능은 얼마나 떨어지느냐가 다음 질문인데, 생각보다 안 떨어진다.

아티피셜 애널리시스의 종합 지능지수에서 V4-플래시는 50점을 받았다. 구글 제미나이 3.6 플래시와 같은 점수다. 메타 뮤즈 스파크 1.1과 Z.ai GLM-5.2보다는 1점 낮다.

다만 상위권과는 벌어진다. 키미 K3가 57점이고, 클로드 오퍼스 5와 페이블 5, GPT-5.6은 V4-플래시보다 9점 이상 높다.

이 지수는 코딩·추론·업무형 과제 등 9개 부문 벤치마크를 묶은 값이다.

값은 105분의 1인데 지능지수는 9점 차. '싼 대신 훨씬 못하다'는 아니라는 뜻이다.

그런데 로이터가 붙인 단서가 중요하다.

표시 단가가 낮아도, 답을 만드는 과정에서 다른 모델보다 추론·출력 토큰을 훨씬 많이 쓰고 단계를 더 거친다면 실제로 나가는 돈은 올라갈 수 있다는 것.

토큰당 값이 아니라 과제 하나를 끝내는 데 드는 값을 봐야 한다는 얘기다. 그 관점에서 보면 테스트 1회 3센트라는 숫자가 오히려 더 중요하다.

그래서 이게 나랑 무슨 상관이냐.

직접 API를 쓰지 않는다면 당장 바뀌는 건 없다. 다만 이런 압박이 쌓이면 값은 내려간다 — 오픈AI가 지난달 30일 GPT-5.6 값을 80%·20% 내린 것도 같은 흐름 위에 있다.

API를 붙여 쓰고 있다면 얘기가 다르다. 지금 쓰는 모델이 그 일에 꼭 필요한 급인지 한 번 재볼 만하다. 번역이나 분류처럼 정답 판정이 쉬운 일은 특히 그렇다.

값이 100배 갈리는데 점수는 9점 차이라면, 질문은 '어느 게 더 좋은가'가 아니라 '내 일에 몇 점이 필요한가'가 된다..

DeepSeek's new model reportedly runs at one one-hundred-and-fifth the cost of Claude.

The benchmark comes from Artificial Analysis, released August 3 local time, reported by Reuters and carried by Dailian.

The model is V4-Flash, the API version DeepSeek released on July 31.

Start with price.

  • $0.14 per million input tokens, $0.28 per million output tokens
  • About 3 cents per test run on average
  • For comparison — Kimi K3 at 86 cents, GPT-5.6 Sol at $1.86, Claude Fable 5 at $3.15

Three cents against $3.15. That is where the 105x comes from.

The next question is how much capability that costs, and the answer is less than you might expect.

On Artificial Analysis's composite intelligence index, V4-Flash scored 50 — the same as Google's Gemini 3.6 Flash, and one point below Meta's Muse Spark 1.1 and Z.ai's GLM-5.2.

The gap to the top is real, though. Kimi K3 scored 57, and Claude Opus 5, Claude Fable 5 and GPT-5.6 all sit more than nine points above V4-Flash.

The index combines benchmarks across nine areas including coding, reasoning and work tasks.

One hundred and fifth the price, nine points of index. Cheap does not mean far worse here.

But Reuters attached a caveat that matters.

A low listed rate can still mean higher real spending if the model burns far more reasoning and output tokens, taking more steps to reach an answer.

What counts is the cost of finishing a task, not the price per token. By that measure the 3-cents-per-run figure is the more useful number.

So what does this mean for you?

If you do not call these APIs directly, nothing changes today. But this pressure pushes prices down — OpenAI's 80% and 20% cuts to GPT-5.6 on July 30 sit on the same curve.

If you do call them, it is worth checking whether the model you are paying for is the tier the job actually needs. That goes double for work where correctness is easy to check, like translation or classification.

When price differs by a hundred times and the score differs by nine points, the question stops being which model is better and becomes how many points your work requires..

Sources · Dailian