xAI가 '그록 봇(Grok Bot)'을 내놨습니다. 아직 얼리 베타입니다.

회사는 이걸 에이전트가 아니라 “진짜 일을 맡길 수 있는 AI 팀원”이라고 불렀습니다. 말장난처럼 들리지만 설계가 실제로 그쪽입니다.

그리고 이 제품이 딛고 선 모델은 며칠 전 프론티어 그룹에 막 올라섰습니다.

■ 봇마다 컴퓨터가 하나씩

에이전트와 팀원은 무엇이 다를까요? 가장 큰 차이가 여기입니다. 봇들은 클라우드에 자기 컴퓨터를 갖고 있습니다.

그래서 사용자가 쓰던 도구에 봇이 직접 로그인해 씁니다. 앱과 받은편지함을 오가며 일하고, 일이 끝날 때까지 붙어 있다가 승인이 필요한 대목에서만 돌아옵니다.

깔끔한 API 나 MCP 가 없는 플랫폼에서도 작동한다고 xAI 는 밝혔습니다. 사람이 화면을 보고 클릭하듯 하기 때문입니다. 사용자가 자리를 비워도 일이 멈추지 않는다는 것이 '자기 컴퓨터'의 실질적인 의미입니다.

xAI 가 든 사내 사례는 셋입니다. 통화 녹취 메모로 CRM 을 갱신하고 후속 메일 초안을 쓰는 영업 봇, 신입 자리 배치를 처리하고 지메일로 받은 청구서를 처리하는 운영 봇, 제품 화면에서 버그를 재현하고 티켓을 등록해 넘기는 엔지니어링 봇.

회사가 이번 발표에서 자랑한 지점은 성능이 아니라 마무리였습니다. 90%와 100% 사이에 큰 차이가 있고, 대부분의 AI 는 거의 다 왔다가 멈추는데 그록 봇은 결과를 실제 도구 안, 사람이 놓을 자리에 놓는다는 것입니다.

■ 봇이 봇에게 일을 넘긴다

봇 하나로 끝일까요? xAI 직원들은 여러 봇을 한꺼번에 돌린다고 했습니다. 그중 하나가 나머지를 관리합니다.

비서실장 격 봇이 위에 있고 그 아래로 받은편지함 정리, 경비, 채용, 버그 수정, 운영 담당이 따로 붙는 식입니다. 봇끼리 메시지를 주고받고 스레드에서 맥락을 공유합니다.

프로젝트가 겹칠 때 사용자가 이 채팅 내용을 저 채팅에 붙여넣지 않아도 같은 건에 대해 말이 맞습니다. 단체 채팅방에 봇들을 넣어두면 자기들끼리 조율합니다. 일을 넘기고 담당을 정하고, 판단이 필요한 대목에서만 사람을 부릅니다.

■ 한 번 옆에서 보게 하면 외운다

설정이 없다는 것을 xAI 는 여러 번 강조했습니다. 다른 도구들은 먼저 워크플로와 루틴을 만들어두라고 하는데, 그록 봇은 그냥 메시지를 보내면 된다는 얘기입니다.

일을 가르치는 방법도 그렇습니다. 다음에 그 일을 할 때 따라다니라고 하면 됩니다. 봇이 단계를 지켜보고 사용자가 어떻게 하는 걸 좋아하는지 기억합니다. 그걸 루틴으로 저장해두고, 고쳐준 것을 반영해 다음부터는 알아서 돌립니다.

다만 후기 하나는 그대로 읽을 필요가 있습니다. 한 사용자는 “워크플로를 한 번 보여줬더니 이제 영원히 돌려도 될 만큼 신뢰한다”며 “내가 검증하고 검토하지 않아도 해내기 때문에 2~3배 효율이 오른 것 같다”고 했습니다.

검증을 건너뛰어서 생긴 효율이라는 뜻이기도 합니다. 회사 공식 페이지에 실린 사용자 후기이지, 검증을 생략해도 된다는 xAI 의 권고는 아닙니다.

■ 며칠 전 올라선 61점

이 제품이 왜 지금 나왔을까요? 밑에 깔린 모델이 막 자리를 옮겼습니다.

xAI 는 대표 모델 그록을 한 달여 만에 4.6 버전으로 올렸습니다. 독립 평가 사이트 AA(Artificial Analysis)가 매긴 그록 4.6의 종합 지능지수는 61점입니다.

그전까지 60점을 넘긴 모델은 앤트로픽 클로드 오퍼스 5와 페이블 5, 오픈AI GPT-5.6 솔, 문샷AI 키미 K3뿐이었습니다. 그록 4.6이 여기 새로 들어갔습니다. GPT-5.6 솔과 동점으로, 클로드 계열 바로 아래 공동 3위입니다. 전작 그록 4.5보다 5점, 그록 4.3보다 23점 높습니다.

가격은 그대로 뒀습니다. 100만 토큰당 입력 2달러, 출력 6달러입니다. 동점인 GPT-5.6 솔은 입력 5달러, 출력 30달러입니다.

같은 점수에 출력 단가는 5분의 1. 이번 발표에서 눈에 띄는 숫자입니다.

■ 마무리를 내세운 이유

벤치마크를 보면 그록 봇이 왜 '마무리'를 앞세웠는지가 드러납니다. 그록 4.6은 특히 에이전틱 작업에서 성적이 올랐습니다.

44개 직종의 실제 업무 결과물 생성 능력을 재는 GDPval-AA v2 에서 62%를 기록해 클로드 오퍼스 5(67%·66%)에 이은 3위, 페이블 5와 동점이었습니다. 그록 4.5는 51%였습니다.

금융 실무를 재는 τ³-뱅킹에서는 51%로 알리바바 큐웬 맥스 3.8과 공동 1위였고, 메신저·이메일·회의록 같은 파편 정보를 조합해 장기 프로젝트를 끝내는지 보는 AA-브리프케이스에서는 Elo 1577로 4위, 페이블 5보다 앞섰습니다.

학습 쪽 설명이 이 대목과 맞물립니다. xAI 는 후속 학습에서 이전 버전 그록 4.5를 프롬프트 최적화에 활용해 SFT 를 수행하고 강화학습을 거쳤다고 밝혔습니다. 그리고 장기 작업을 할 때 모델이 스스로 결과물의 오류를 검증하도록 동작 성향을 개선했다고 했습니다.

그러면 벤치마크 점수가 실제 마무리를 보장할까요? 그록 봇이 내세운 '90%에서 멈추지 않는다'는 주장의 근거는 여기에 있습니다. 다만 벤치마크 점수와 실제 업무에서의 마무리는 같은 것이 아닙니다.

■ 지금 누가 쓸 수 있나

베타이고, 발표일부터 슈퍼그록 헤비와 커서 울트라, 커서 팀스 프리미엄 구독자가 쓸 수 있습니다. 데스크톱과 iOS 이고, 발표 페이지의 내려받기 버튼은 macOS 용입니다.

기업 사용자는 대기자 명단에 이름을 올리는 단계입니다. 한국어 지원 여부나 국내 출시 시점은 이번 발표에 없습니다.

원래는 사내 프로토타입이었다고 합니다. 영업, 마케팅 캠페인, 사무 운영, 버그 수정까지 회사 전체로 번져서 밖에도 열게 됐다는 것이 xAI 의 설명입니다.

모델은 프론티어에 올랐고 값은 그대로이며 제품은 베타입니다. 셋 중 확인된 것은 앞의 둘뿐입니다.

xAI has released Grok Bot. It is still an early beta.

The company calls it not an agent but “AI teammates you can actually hand work to”. It sounds like wordplay, but the design really does point that way.

And the model underneath it had just climbed into the frontier tier days earlier.

Each bot gets its own computer

What separates an agent from a teammate? This is the main difference. The bots have their own computers in the cloud.

So a bot logs into the tools you already use and works them directly. It moves between apps and the inbox, stays on the job until it is done, and comes back only where approval is needed.

xAI says it works even on platforms with no clean API or MCP, because it looks at the screen and clicks the way a person would. That a job does not stop when you step away is the practical meaning of 'its own computer'.

xAI gave three in-house examples: a sales bot that updates the CRM from call notes and drafts follow-ups; an operations bot that handles new-hire desk assignments and invoices arriving by Gmail; and an engineering bot that reproduces bugs in the product and files tickets.

What the company boasted about was not capability but finishing. There is a large gap between 90% and 100%, it argues, and most AI stops just short — whereas Grok Bot puts the result inside the actual tool, where a person would have put it.

Bots hand work to bots

Is one bot the end of it? xAI staff run several bots at once, the company says, with one of them managing the rest.

A chief-of-staff bot sits on top, with separate bots for inbox triage, expenses, recruiting, bug fixes and operations below it. The bots message each other and share context in threads.

When projects overlap, you do not have to paste one chat into another for them to stay consistent. Put the bots in a group chat and they coordinate among themselves — handing off work, assigning owners, and calling a human only where judgement is needed.

Show it once and it remembers

xAI stressed repeatedly that there is no setup. Other tools ask you to build workflows and routines first; with Grok Bot you just send a message.

Teaching it works the same way. Tell it to follow along next time you do the task. The bot watches the steps, remembers how you like things done, saves it as a routine, folds in your corrections, and runs it itself from then on.

One testimonial deserves reading as written. A user said that after showing it a workflow once, “I trust it enough to run forever,” and that efficiency was up “two or three times because it gets there without me verifying and reviewing.”

Which also means the efficiency came from skipping verification. It is a user testimonial on the company's own page, not xAI advising anyone to skip checks.

The 61 points it stands on

Why now? Because the model underneath just moved.

xAI updated Grok to version 4.6 a little over a month after the previous release. On the independent evaluation site Artificial Analysis, Grok 4.6 scores 61 on the composite intelligence index.

Until then only Anthropic's Claude Opus 5 and Fable 5, OpenAI's GPT-5.6 Sol and Moonshot's Kimi K3 had passed 60. Grok 4.6 joins them, level with GPT-5.6 Sol and jointly third behind the Claude line. It is 5 points above Grok 4.5 and 23 above Grok 4.3.

The price did not move: $2 per million input tokens and $6 per million output. GPT-5.6 Sol, on the same score, is $5 and $30.

Same score, a fifth of the output price. That is the number that stands out.

Why finishing was the pitch

The benchmarks show why Grok Bot led with finishing. Grok 4.6 improved most on agentic work.

On GDPval-AA v2, which measures producing real work output across 44 occupations, it scored 62% — third behind Claude Opus 5 (67% and 66%) and level with Fable 5. Grok 4.5 scored 51%.

On τ³-Banking, which tests practical financial work, it tied for first at 51% with Alibaba's Qwen Max 3.8. On AA-Briefcase, which tests completing long projects by assembling fragmentary information from chats, email and meeting notes, it placed fourth at Elo 1577, ahead of Fable 5.

The training notes line up with this. xAI says post-training used the earlier Grok 4.5 in a prompt-optimisation workflow for supervised fine-tuning, followed by reinforcement learning — and that the model's behaviour was tuned to check its own work for errors on long-horizon tasks.

So does a benchmark score guarantee the finish? That is where the 'we don't stop at 90%' claim comes from. A benchmark score and finishing a real job are still not the same thing.

Who can use it now

It is a beta. From launch day it is available to SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium subscribers, on desktop and iOS; the download button on the announcement page is for macOS.

Enterprise users are at the waitlist stage. Korean-language support and a Korean launch date do not appear in the announcement.

It began as an internal prototype. It spread across sales, marketing campaigns, office operations and bug fixing inside the company, which is why xAI opened it up.

The model reached the frontier, the price held, and the product is in beta. Only the first two are established.

Sources · xAI News · TheElec