AI가 변호사시험 선택형 150문항을 다 맞혔다.

리걸테크 회사 쥬리서포트가 7월 31일 밝힌 결과다. 제15회 변호사시험 선택형 전 과목이 대상이었다.

그런데 이 실험에서 눈여겨볼 건 만점이 아니다.

같은 AI에서 도구 하나만 뗐더니 결과가 어떻게 달라졌는지다.

먼저 구성부터. 과목마다 다른 모델을 썼다.

  • 공법 40문항 — GPT-5.6 Sol
  • 민사법 70문항 · 형사법 40문항 — Claude Opus 5

각 모델에는 법령·판례·교재를 검색하는 자체 리서치 도구를 붙였다. 과목별로 세 차례씩, 모두 아홉 번 시험을 봤다.

그리고 대조군을 뒀다. 같은 모델에서 리서치 기능만 뗀 것이다.

대조군은 아홉 번 중 단 한 번도 만점을 못 받았다.

  • 공법 — 평균 33점 (40점 만점)
  • 민사법 — 세 차례 모두 65점 (70점 만점)
  • 형사법 — 37점 · 35점 · 36점 (40점 만점)

모델은 같았다. 법령·판례를 찾아보는 절차를 붙였느냐 아니냐만 달랐다.

민사법이 특히 눈에 걸린다. 세 번 다 65점이다.

매번 다른 문제를 틀린 게 아니라 같은 데서 걸렸을 가능성이 높다는 뜻인데, 도구를 붙이자 그게 전부 교정됐다.

하희봉 쥬리서포트 대표는 "AI 단독의 성능이 아니라, 법령·판례·교재를 실제로 찾아 확인하는 리서치 절차를 붙였을 때 무엇이 달라지는지를 측정한 실험"이라고 했다.

회사는 한계도 같이 밝혔다. 이 대목은 그대로 옮긴다.

이번 결과는 5지선다 선택형에 한정되고, 사례형·기록형 답안을 쓰는 능력을 뜻하지 않는다. 실제 시험의 시간 제한 조건과도 다르다.

변호사 일의 핵심은 선택형이 아니라 서면이다. 그러니 "AI가 변호사시험에 붙었다"고 읽으면 과한 해석이다.

데이터도 공개 자료만 썼다고 한다. 공개 법령·판례와 자체 보유 법률 교재를 DB로 만들어 검색하게 했고, 비공개 자료는 쓰지 않았다는 설명이다.

그리고 실험에 쓴 리서치 도구를 오픈소스로 풀었다. '쥬리서포트 플러그인'과 법률서적을 DB화한 'legal-books'를 MIT 라이선스로 배포했다.

그래서 이게 나랑 무슨 상관이냐.

법을 다루는 사람이 아니어도 가져갈 게 있다.

AI한테 그냥 물으면 그럴듯하게 틀린 답이 온다. 근거를 찾아보게 만들면 달라진다 — 이 실험이 그걸 점수로 보여줬다.

챗봇에 검색을 켜거나, 원문을 붙여넣고 "여기서만 찾아서 답해"라고 시키는 것도 같은 원리다.

AI가 똑똑해진 게 아니라, 찾아보게 시켰더니 달라진 거다..

An AI answered all 150 multiple-choice questions on Korea's bar exam correctly.

JuriSupport, a legal-tech company, announced the result on July 31. It covered every multiple-choice subject on the 15th Korean bar examination.

But the perfect score is not the interesting part.

What matters is what happened when they removed one tool from the same AI.

First, the setup. Different models handled different subjects.

  • Public law, 40 questions — GPT-5.6 Sol
  • Civil law (70) and criminal law (40) — Claude Opus 5

Each model was connected to an in-house research tool that searches statutes, case law, and textbooks. Each subject was run three times, nine sittings in total.

Then came the control group: the same models with the research function stripped out.

The control group never once scored full marks across those nine runs.

  • Public law — 33 out of 40 on average
  • Civil law — 65 out of 70, all three times
  • Criminal law — 37, 35 and 36 out of 40

The model was the same. The only difference was whether it had to go look the law up.

Civil law stands out: 65 on all three attempts.

That suggests it was failing on the same points rather than missing different questions each time. Attaching the tool corrected all of them.

Ha Hee-bong, the CEO of JuriSupport and a practicing lawyer, described it as "an experiment measuring what changes when you attach the process of actually looking up statutes, precedents and textbooks, rather than testing the AI's ability on its own."

The company also stated the limits, which are worth repeating as given.

The result applies only to five-option multiple choice. It says nothing about writing case-based or record-based answers, and the conditions differed from the real exam's time limits.

The core of legal work is written argument, not multiple choice. Reading this as "AI passed the bar" overstates it.

The data was public, the company says. It built a searchable database from published statutes and case law plus its own legal textbooks, and used no private material.

It also open-sourced the research tooling: the JuriSupport plugin and legal-books, a database of legal texts, both under the MIT license.

So what does this mean for you?

There is something here even if you never touch the law.

Ask an AI cold and you get confident, wrong answers. Make it go find the source and things change — this experiment put a score on that gap.

Turning on search in a chatbot, or pasting the source text and saying "answer only from this," works on the same principle.

The AI did not get smarter. It was made to look things up..

Sources · Edaily