"당신 밑에 AI 에이전트 6개를 두고 매니저처럼 일하라." 마이크로소프트가 내놨다는 이 가이드가 요즘 직장인들 사이에서 회자됩니다.

SBS가 직장인 앱 블라인드의 문성욱 대표를 인터뷰한 내용입니다. 매니저였든 아니었든 어떤 자리에 있든 AI를 써서 일해야 하고, 그 결과물을 검수하고 정리하는 역할까지 맡게 됐다는 이야기였습니다.

그런데 같은 인터뷰에서 정작 사람들이 스트레스를 받는 지점은 따로 나옵니다. 에이전트를 여럿 굴리는 게 아니라, 그것들이 뱉어낸 저품질 결과물을 검수하고 고치는 일이 늘었다는 것입니다.

■ 정작 늘어난 건 검수 시간

인터뷰는 이 현상을 '슬롭'이라고 불렀습니다. 생성형 AI가 대량으로 찍어내는 무의미하거나 품질 낮은 작업물을 비판적으로 가리키는 말입니다.

AI로 일을 더 빨리, 더 많이 할 수 있는 건 사실입니다. 다만 그만큼 회사의 기대치도 올라갑니다. 여기에 검수해서 고쳐야 할 일까지 늘어나면서 스트레스를 받는다는 설명이었습니다.

그렇다면 에이전트를 몇 개 굴리느냐는 애초에 문제의 핵심이 아닐 수 있습니다. 검수 시간을 줄이는 쪽으로 일을 쪼개는 게 먼저입니다.

인터뷰는 실무에 깊숙이 쓰이기까지 늦어도 2~3년으로 봤습니다. 규모가 큰 회사일수록 타격이 크고, 작은 회사는 한 사람의 능력을 키우는 기회가 될 수 있다는 전망도 덧붙였습니다.

■ 국내에서 이미 갈라진 자리

어떤 도구를 쓰느냐부터 직군에 따라 갈렸습니다.

지난해까지는 챗GPT가 압도적이었습니다. 한국 직장인의 70% 이상이 챗GPT를 쓴다는 이야기가 나왔습니다. 그런데 미국은 올해 1월부터, 한국은 올해 중반부터 양상이 뒤집혔다는 게 인터뷰의 설명입니다.

코딩이나 소프트웨어 엔지니어링을 조금이라도 하는 회사에서는 클로드 사용률이 앞섰습니다. 현대차와 네이버, 비바리퍼블리카, NC소프트, LG전자 같은 곳이 예로 꼽혔습니다. 반면 공무원과 경찰청, 기업은행, 철도공사, 한수원 쪽에서는 챗GPT 이용이 우세했습니다.

제미나이는 지메일과 구글 드라이브처럼 구글 생태계 안에서는 앞서지만, 소프트웨어 엔지니어링으로 내려오면 다른 모델보다 떨어진다는 평가를 받는다고 했습니다.

보안 때문에 자체 AI를 쓰라는 회사도 많은데, 품질이 외부 모델에 못 미쳐 불만이 쌓인다는 이야기도 나왔습니다. SK그룹과 LG그룹이 예로 들렸고, 한국만의 현상은 아니라고 했습니다. 아마존과 마이크로소프트조차 클로드를 더 많이 쓴다는 조사 결과를 함께 전했습니다.

■ 일을 쪼개는 기준은 '호출 횟수'다

그럼 무엇을 기준으로 일을 나눠 맡겨야 할까요?

앤트로픽 도움말에 기준이 적혀 있는데, 주제가 아니라 '몇 번 찾아봐야 하는 일이냐'로 나눕니다. 이게 검수 시간과 직결됩니다.

웹 검색은 한두 번의 도구 호출로 답이 나오는 사실 질의에 맞습니다. 날씨를 확인하거나 특정 회사 정보를 얻거나 최신 뉴스 헤드라인을 가져오는 일입니다. 무엇을 찾을지는 대략 알려주되 검색과 정리는 맡기고 싶을 때도 여기에 해당합니다.

확장 사고는 최신 정보가 필요 없는 복잡한 추론에 맞습니다. 수학 문제를 풀거나 코드를 디버깅하거나 개념을 분석하는 일입니다. 검색이나 연동 없이 그냥 더 깊이 생각하게 하고 싶을 때 씁니다.

리서치는 도구 호출이 다섯 번 이상 필요하고 1~3분이 걸리는 일에 맞습니다. 웹과 연동된 자료를 여러 출처에서 종합해 보고서를 만드는 작업입니다. 앤트로픽은 계획을 세우고 실행까지 하려면 확장 사고와 리서치를 함께 쓰라고 권합니다.

기준이 이렇게 잡히면 쪼개는 방법도 정해집니다. 한두 번이면 되는 일을 리서치에 던지면 시간만 쓰고, 다섯 번 넘게 파야 할 일을 한 번 물어보면 얕은 답이 나와 검수 부담으로 돌아옵니다.

■ 리서치를 켜는 법과 한도

리서치는 유료 요금제에서만 됩니다. 프로·맥스·팀·엔터프라이즈이고, 웹과 데스크톱, 모바일에서 쓸 수 있습니다.

채팅창 왼쪽 아래 '+' 버튼을 누르고 'Research'를 고르면 창 아래에 파란 표시가 나타납니다. 다시 누르면 꺼집니다. 한 가지 전제가 있습니다. 웹 검색이 켜져 있어야 리서치가 작동합니다.

켰는데도 리서치를 안 하는 것 같으면 어떻게 할까요?

앤트로픽은 직접 시키라고 안내합니다. "리서치 도구를 써서…"라고 말로 지시하라는 것입니다. 지메일이나 구글 문서 같은 내부 자료를 연결해 뒀는데 안 끌어오면 "[해당 자료]에서 관련 맥락을 가져와"라고 짚어 주라고 했습니다.

비용 쪽 주의도 붙어 있습니다. 리서치는 일반 대화와 같은 한도를 쓰지만, 여러 출처를 가져오고 긴 답을 내놓기 때문에 한도를 더 빨리 소진한다고 명시했습니다. 에이전트를 여럿 굴린다는 건 한도도 그만큼 빨리 쓴다는 뜻입니다.

■ 동시에 굴리려면 어디서 도는지 본다

여러 작업을 겹쳐 돌릴 생각이라면 어디서 실행되는지가 중요해집니다.

앤트로픽은 코워크가 이제 그냥 클로드로 합쳐지고 있다고 안내했습니다. 필요한 걸 말하면 빠른 답으로 끝낼지 작업으로 돌릴지 클로드가 판단한다는 것이고, 프로와 맥스에 순차 적용 중입니다.

클라우드에서 돌면 노트북을 닫아도 작업이 이어집니다. 예약 작업도 클라우드에서 돌아 컴퓨터가 켜져 있을 필요가 없습니다. 세션과 파일은 계정에 붙어 있어 데스크톱과 웹, 모바일에서 같은 것을 봅니다.

다만 전부 클라우드로 되는 건 아닙니다. 로컬 파일 접근과 로컬 커넥터, 브라우저 사용은 데스크톱 앱을 거칩니다. 클라우드 세션이 내 컴퓨터 폴더를 읽고 쓰려면 데스크톱 앱이 켜져 있어야 합니다. 로컬 MCP 서버를 포함한 커넥터와 플러그인도 데스크톱에서만 동작합니다.

■ 검수를 줄이는 쪽으로

정리하면 이렇습니다.

  • 몇 번 찾아봐야 하는 일인지 먼저 셉니다 — 한두 번이면 웹 검색, 다섯 번 이상이면 리서치입니다
  • 찾을 게 없는 일은 검색을 끕니다 — 계산과 디버깅은 확장 사고 쪽입니다
  • 안 돌면 말로 직접 시킵니다 — 앤트로픽이 안내하는 방법이 그것입니다
  • 내부 자료는 어디서 가져올지 짚어 줍니다 — 연결만 해 두면 알아서 오지 않을 수 있습니다
  • 한도 소진 속도를 감안합니다 — 리서치는 같은 한도를 더 빨리 씁니다
  • 로컬 파일이 걸린 작업은 데스크톱에 남깁니다 — 클라우드만으로는 안 됩니다
  • 검수할 사람이 하나라는 걸 잊지 않습니다 — 동시에 돌린 만큼 볼 것도 늘어납니다

마지막 항목이 이 글의 요점입니다. 에이전트를 여섯 개 붙여 준다는 말은 일이 6분의 1로 줄어든다는 뜻이 아니라, 검수해야 할 결과물이 여섯 갈래로 들어온다는 뜻이기도 합니다. 늘리기 전에 각각이 무엇을 맡는지부터 정해 두는 편이 낫습니다.

"Put six AI agents under you and work like a manager." A guideline attributed to Microsoft has been making the rounds among office workers here.

It comes from an SBS interview with Moon Sung-wook, head of the workplace app Blind. Whether or not you were ever a manager, the point was, everyone now has to work through AI - and take on reviewing and cleaning up what it produces.

Yet the same interview locates the actual stress somewhere else: not in running several agents, but in the growing work of checking and fixing the low-quality output they produce.

■ What actually grew was review time

The interview calls this "slop" - a critical term for the meaningless or low-quality output generative AI produces in bulk.

AI does let people work faster and do more. But expectations rise with it, and when the work of reviewing and correcting piles on top, the result is stress rather than relief.

If that is the case, how many agents you run was never the core question. Splitting work so that review time shrinks comes first.

The interview put deep practical adoption at two to three years out at the latest, with larger companies taking more of the hit and smaller ones getting a chance to raise what one person can do.

■ Where Korean workplaces already split

Even the choice of tool has divided along job lines.

Through last year ChatGPT dominated - more than 70% of Korean office workers were said to use it. That flipped from January this year in the US and from mid-year in Korea, according to the interview.

At companies doing any coding or software engineering, Claude took the larger share, with Hyundai Motor, Naver, Viva Republica, NCSoft and LG Electronics named as examples. In the civil service, the police agency, IBK, Korea Railroad and KHNP, ChatGPT use stayed dominant.

Gemini leads inside Google's own ecosystem - Gmail, Google Drive - but is rated weaker than rivals once the work descends to software engineering.

Many companies mandate in-house AI for security reasons, and complaints accumulate because quality trails outside models. SK Group and LG Group were cited, and the interview noted this is not only a Korean pattern: even Amazon and Microsoft use Claude more, per its survey.

■ The split is by number of tool calls

So what should you divide the work by?

Anthropic's help pages give a criterion that is not topic but how many lookups a job needs - and that maps directly onto review time.

Web search fits straightforward factual queries answerable in one or two tool calls: checking the weather, getting details on a company, pulling recent news headlines. It also fits when you want to say roughly what to search and leave the searching and analysis to Claude.

Extended thinking fits complex reasoning that needs no recent information - solving mathematical problems, debugging code, analyzing concepts. Use it when you do not want searching or integrations, just harder thinking.

Research fits work needing five or more tool calls over one to three minutes, synthesizing multiple sources across the web and your integrations into an in-depth report. Anthropic recommends combining extended thinking with research when you want planning and execution together.

With that criterion the splitting rule follows. Throw a two-call job at research and you only spend time; ask a five-call job in a single question and the shallow answer comes back as review work.

■ Turning research on, and what it costs

Research requires a paid plan - Pro, Max, Team or Enterprise - on web, Desktop or Mobile.

Click the "+" button at the bottom left of the chat and select Research; a blue indicator appears at the bottom of the window, and clicking it again disables it. One prerequisite: web search must be turned on for research to function.

What if you enable it and Claude does not seem to be researching?

Anthropic says to steer it directly - tell Claude to use the research tool. If internal sources like Gmail or Google Docs are connected but not being pulled, it advises prompting Claude to pull relevant context from that source by name.

There is a cost note too. Research is subject to the same limits as standard conversations, but sessions can use up those limits faster because Claude retrieves multiple sources and produces comprehensive responses. Running several agents means burning through limits faster as well.

■ If you run them in parallel, check where they run

Once you plan to overlap tasks, where execution happens starts to matter.

Anthropic notes Cowork is now folding into Claude itself: ask for what you need and Claude decides whether it is a quick answer or a task, rolling out gradually to Pro and Max plans.

Work in the cloud continues after you close your laptop, and scheduled tasks run in the cloud so your computer no longer needs to be awake. Sessions and files live with your account, so desktop, web and mobile show the same things.

Not everything moves to the cloud, though. Local file access, local connectors and browser use go through the Desktop app - a cloud session can read and write folders on your computer only while the desktop app is running. Connectors and plugins that include local MCP servers work on desktop only.

■ Optimize for less reviewing

In short:

  • Count the lookups first - one or two means web search, five or more means research
  • Turn search off when there is nothing to look up - calculation and debugging belong to extended thinking
  • If it does not engage, tell it directly - that is Anthropic's own advice
  • Name the internal source to pull from - connecting it does not guarantee it gets used
  • Account for limit burn - research spends the same allowance faster
  • Leave local-file work on desktop - the cloud alone will not do it
  • Remember there is still one reviewer - everything you run in parallel comes back to you

That last point is the whole article. Six agents under you does not mean the work drops to a sixth; it can also mean output arrives in six streams, all of it needing review. Deciding what each one owns is worth doing before adding more.