우버는 지난해 엔지니어 5,000여 명에게 AI 코딩 도구를 적극적으로 쓰라고 권했습니다. 생산성을 끌어올리기 위해서였습니다.
그리고 넉 달 뒤, 회사는 그해 배정한 AI 예산이 이미 바닥났다는 사실을 확인했습니다. 1년을 버티라고 잡아둔 돈이었습니다.
우버가 내린 결정은 도구를 더 사주는 것이 아니었습니다. 직원 한 사람당 월 사용 한도를 약 1,500달러로 묶었습니다.
AI를 쓰지 못해서가 아니라, AI가 비싸서 손을 놓는 일이 벌어지고 있습니다. 그리고 이건 우버 한 곳의 이야기가 아닙니다.
■ 2028년, 5곳 중 1곳이 AI를 접는다
글로벌 시장조사업체 가트너는 2028년까지 전 세계 조직의 20%가 AI 관련 비용을 통제하지 못해 업무 적용을 포기할 것으로 내다봤습니다. 다섯 곳 중 한 곳입니다.
포기한 자리를 무엇이 대신할까요. 가트너는 이들이 기존 소프트웨어 개발 체제로 되돌아갈 것이라고 봤습니다. 투자 대비 효과가 낮은 업무부터 AI 적용을 줄이거나 옛 시스템으로 교체하는 흐름입니다.
기술이 모자라서가 아니라 돈이 모자라서 뒤로 가는 셈입니다.
■ "예산을 넘겼다" 93%…예외가 아니라 거의 전부
우버의 사례가 유별난 것인지 확인할 수 있는 숫자가 있습니다. 맥킨지가 지난 5월 주요 5개 산업의 기업 75곳을 대상으로 AI 핀옵스(FinOps) 조사를 진행했습니다.
응답 기업의 93%가 AI 예산을 초과 집행했다고 답했습니다. 예산을 지킨 곳이 열에 하나도 되지 않았다는 뜻입니다.
이름이 알려진 기업들도 예외가 아니었습니다. 마이크로소프트는 올해 5월 수천 명의 내부 직원에게 제공하던 클로드 코드 라이선스를 중단하고, 자사 깃허브 코파일럿으로 전환했습니다.
■ 왜 에이전트는 유독 비싼가
비용이 튀는 지점은 생각보다 구체적입니다. 스스로 작업을 수행하는 에이전트형 AI가 지목됐습니다.
분석에 따르면 에이전트형 AI가 쓰는 전체 비용 가운데 약 60%는 처음 답을 만들어내는 과정이 아니라, 그 결과물을 수정하고 보완하는 반복 작업에서 발생합니다. 답을 내는 것보다 고치는 데 돈이 더 드는 구조입니다.
토큰 소비량으로 보면 차이는 더 벌어집니다. 에이전트형 AI는 같은 업무를 수행할 때 일반 챗봇보다 최대 30배 많은 토큰을 쓸 수 있습니다. 여기서 결과가 기준을 충족하지 못해 작업을 다시 수행하면, 전체 맥락을 처음부터 다시 처리하면서 토큰 소비가 최대 50배까지 뛰기도 합니다.
한 번에 끝나지 않을수록 청구서가 기하급수적으로 불어난다는 얘기입니다.
■ 정액제에서 종량제로…쓸수록 예측이 어려워진다
시장의 가격 구조 자체도 바뀌고 있습니다. 주요 개발사들이 시장을 선점하려고 무료나 저가로 서비스를 뿌리던 초기와 달리, 최근 대규모언어모델(LLM) 시장은 소수 업체 중심으로 재편되는 양상입니다.
요금 체계는 정액제에서 사용량에 따라 비용이 늘어나는 종량제로 빠르게 옮겨가고 있습니다. 기업 입장에서는 AI 사용량이 늘어날수록 최종 비용을 예측하고 관리하기가 어려워지는 구조입니다.
그렇다면 기다리면 싸질까요. 가트너는 이런 시장 변화를 근거로, 앞으로 AI 이용 비용이 지금보다 낮아지기는 어려울 것으로 내다봤습니다.
■ 가트너가 내놓은 답은 '관리'
가트너가 제시한 대응책은 'AI 핀옵스'입니다. AI 사용량과 지출, 성과를 지속적으로 추적·관리해 불필요한 비용을 줄이고 투자 효율을 높이는 방식입니다.
구체적으로는 업무별로 적합한 AI 모델을 선택하고, 필요 이상으로 긴 답변이나 연산을 줄이라는 것입니다. 가장 좋은 모델을 모든 일에 물리는 것이 답이 아니라는 뜻이기도 합니다.
AI 업계 관계자는 "그동안에는 AI를 얼마나 빠르고 폭넓게 도입하느냐가 경쟁력으로 여겨졌지만 비용 부담이 본격화하면서 상황이 달라지고 있다"며 "앞으로는 실질적인 생산성 향상이 가능한 업무를 선별해 AI를 적용하는 전략이 중요해질 것"이라고 말했습니다.
■ 남는 질문은 '누가 계속 쓸 수 있는가'
비용이 경쟁의 기준이 되면 따라오는 문제가 하나 더 있습니다. 고성능 프론티어 모델의 높은 비용 구조가 지속될 경우, AI 활용이 충분한 자본을 갖춘 대기업이나 일부 산업에 집중될 가능성이 있다는 우려입니다.
AI가 격차를 줄이는 도구가 될 것이라는 기대와는 정반대 방향입니다.
지금까지 우리는 AI를 얼마나 빨리 들여오느냐를 놓고 경쟁해 왔습니다. 앞으로는 얼마나 오래 감당할 수 있느냐가 그 자리를 대신할지도 모릅니다.
여러분의 조직은 지난달 AI에 얼마를 썼는지 알고 계십니까?
Last year Uber urged some 5,000 engineers to make active use of AI coding tools, aiming to lift productivity.
Four months later the company found that the AI budget set aside for the year was already gone. That money was meant to last twelve months.
Uber's response was not to buy more. It capped usage at roughly $1,500 per person per month.
Companies are not stepping back from AI because they cannot use it, but because it costs too much. And Uber is not alone.
One in five will drop AI by 2028
Gartner expects 20% of organizations worldwide to abandon applying AI to their work by 2028, unable to control the associated cost. One in five.
What replaces it? Gartner expects them to fall back to conventional software development, scaling back AI first on work with weak returns or swapping it for legacy systems.
Going backwards not for lack of technology, but for lack of money.
93% went over budget — hardly the exception
There is a number that shows whether Uber was unusual. In May, McKinsey surveyed 75 companies across five major industries on AI FinOps.
93% of respondents said they had overrun their AI budget. Fewer than one in ten stayed inside it.
Well-known companies were no exception. In May, Microsoft cancelled the Claude Code licences it had provided to thousands of internal staff and moved them to its own GitHub Copilot.
Why agents cost so much more
The place the money leaks is fairly specific: agentic AI that carries out tasks on its own.
By the analysis, about 60% of an agent's total cost comes not from producing the first answer but from the repeated work of correcting and refining that output. Fixing costs more than answering.
In tokens the gap widens further. An agent can consume up to 30 times more tokens than an ordinary chatbot for the same task. And when a result falls short and the work is redone, reprocessing the entire context can push token use up to 50 times.
The less often it succeeds in one pass, the faster the bill compounds.
From flat rates to metered — harder to predict the more you use
The pricing structure is shifting too. Unlike the early days when vendors gave the service away or priced it low to win the market, the large language model market is now consolidating around a few players.
Pricing is moving quickly from flat rates to metered billing. The more a company uses AI, the harder its final bill becomes to forecast and control.
So will waiting make it cheaper? On the strength of those market shifts, Gartner does not expect AI to cost less than it does today.
Gartner's answer is management
The proposed remedy is 'AI FinOps' — continuously tracking AI usage, spend and results to cut waste and improve return.
In practice: pick the model suited to each task, and trim answers or computation that run longer than needed. Pointing the best available model at everything is not the answer.
An industry official said that "competitiveness used to be judged by how fast and how broadly you adopted AI, but that is changing as the cost burden becomes real," adding that "the strategy that matters now is selecting the work where AI delivers real productivity gains."
The open question is who can keep paying
When cost becomes the basis of competition, another problem follows. If the cost structure of high-end frontier models holds, AI use may concentrate in large, well-capitalised companies and a handful of industries.
That runs directly against the expectation that AI would narrow gaps.
Until now the contest has been about how quickly you could bring AI in. From here it may be about how long you can afford to keep it.
Do you know how much your organization spent on AI last month?