법률 리서치 도구를 결합한 생성형 AI 가 변호사시험 선택형 전 과목에서 만점을 받았습니다.
리걸테크 기업 쥬리서포트가 7월 31일 밝힌 결과입니다. 대상은 제15회 변호사시험 선택형 150문항 전부였습니다.
눈여겨볼 것은 만점 자체가 아닙니다. 같은 AI 모델에서 리서치 도구만 떼어낸 대조군은 한 차례도 만점을 기록하지 못했습니다.
■ 과목마다 다른 모델
어떤 모델을 썼을까요? 과목별로 갈렸습니다. 공법 40문항은 GPT-5.6 Sol, 민사법 70문항과 형사법 40문항은 클로드 오푸스 5 였습니다.
각 모델에는 법령과 판례, 교재를 검색할 수 있는 자체 법률 리서치 도구를 연결했습니다. 과목별로 세 차례씩 모두 아홉 번 시험을 진행했고, 아홉 번 전부 만점이었습니다. 채점은 법무부가 공개한 정답을 기준으로 이뤄졌습니다.
■ 도구를 떼자 무너졌다
대조군은 동일한 AI 모델에서 법률 리서치 기능만 제거한 조건이었습니다.
공법 평균 33점(40점 만점), 민사법 세 차례 모두 65점(70점 만점), 형사법 37점과 35점, 36점(40점 만점). 만점은 한 번도 나오지 않았습니다.
쥬리서포트는 이번 결과가 AI 자체의 성능보다 법령과 판례를 실시간으로 검색하고 검증하는 리서치 기능의 효과를 보여준 사례라고 설명했습니다.
■ 비공개 자료는 쓰지 않았다
그러면 무엇을 검색하게 했을까요? 공개된 법령·판례 데이터와 회사가 자체 보유한 법률 교재였습니다. 이를 데이터베이스로 만들어 검색하도록 구성했고, 별도의 비공개 자료는 사용하지 않았다고 밝혔습니다.
실험에 사용한 법률 리서치 도구도 오픈소스로 공개했습니다. 사건 관리와 서면 작성 등을 지원하는 쥬리서포트 플러그인, 그리고 법률서적을 데이터베이스화한 legal-books 를 MIT 라이선스로 배포해 누구나 무료로 쓸 수 있게 했습니다.
■ 대표가 짚은 지점
하희봉 쥬리서포트 대표는 변호사입니다. 그는 “AI 단독의 성능이 아니라, 법령·판례·교재를 실제로 찾아 확인하는 리서치 절차를 붙였을 때 무엇이 달라지는지를 측정한 실험”이라고 설명했습니다.
이어 도구 없는 조건에서 반복적으로 틀린 문항이 도구를 붙이자 전부 교정됐다는 점을 들어, 변호사 업무에서 AI 를 쓸 때 검증 가능한 리서치 인프라가 왜 필수인지 보여준다고 분석했습니다.
■ 만점이 말하지 않는 것
그러면 AI 가 변호사시험을 통과했다고 말할 수 있을까요? 회사 측이 먼저 선을 그었습니다.
이번 결과는 5지선다 선택형에 한정되며 사례형과 기록형 답안 작성 능력을 의미하지 않는다는 것입니다. 실제 시험의 시간 제한 조건과도 다르다는 점을 덧붙였습니다.
선택형 150문항 만점, 그러나 측정하지 않은 사례형. 이번 발표가 남긴 기록은 모델의 성능보다 절차에 관한 것입니다.
A generative AI paired with legal research tools scored full marks across every multiple-choice subject of Korea’s bar examination.
The legaltech company JuriSupport announced the result on 31 July. The target was all 150 multiple-choice questions of the 15th bar examination.
The perfect score is not the striking part. A control group using the same AI models with only the research tools removed never once reached full marks.
A different model per subject
Which models were used? It varied by subject. GPT-5.6 Sol took the 40 public-law questions; Claude Opus 5 took the 70 civil-law and 40 criminal-law questions.
Each model was connected to the company’s own legal research tool, able to search statutes, precedents and textbooks. Three runs per subject, nine in total, and all nine were perfect. Marking followed the answer key published by the Ministry of Justice.
Remove the tools and it collapsed
The control condition was the same AI models with only the legal research function stripped out.
Public law averaged 33 of 40; civil law scored 65 of 70 on all three runs; criminal law came in at 37, 35 and 36 out of 40. Not one perfect score.
JuriSupport said the outcome demonstrates the effect of research that searches and verifies statutes and precedents in real time, rather than the raw capability of the model.
No private material
So what was it searching? Publicly available statute and precedent data, plus legal textbooks held by the company. These were built into a database for retrieval, and the company said no separate non-public material was used.
The research tooling itself was open-sourced. The JuriSupport plugin, which supports case management and document drafting, and legal-books, a database of legal texts, were released under the MIT licence for anyone to use free of charge.
What the chief executive stressed
Ha Hee-bong, the chief executive of JuriSupport, is a practising lawyer. He described it as “an experiment measuring what changes when you attach a research procedure that actually looks up and verifies statutes, precedents and textbooks, rather than the performance of the AI on its own”.
He added that items missed repeatedly without the tools were all corrected once the tools were attached, which he said shows why verifiable research infrastructure is essential when AI is used in legal work.
What a perfect score does not say
So can we say an AI passed the bar exam? The company drew the line itself.
The result is confined to five-option multiple choice and does not indicate the ability to write case-based or record-based answers. It also noted the conditions differ from the real exam’s time limits.
A perfect 150 on the multiple-choice paper, and case-based writing left unmeasured. What this announcement records is a point about procedure rather than model capability.
Sources · Edaily