오픈AI가 아직 내놓지 않은 모델로 수학 난제 10건을 풀었다고 밝혔다.

모델 이름은 '아스트라'. 수학과 이론컴퓨터과학에서 10년 넘게 안 풀리던 문제들이다.

사실 이런 발표는 처음이 아니다. 그래서 이번 건 다른 데를 봐야 한다.

작년에 같은 주장을 했다가 반박당했다

오픈AI의 당시 과학 담당 부사장 케빈 웨일은 GPT-5가 에르되시 문제 10건을 풀었다고 했었다.

에르되시 문제 데이터베이스를 관리하는 토머스 블룸이 이를 "심각한 왜곡"이라고 반박했다. 모델이 문제를 푼 게 아니라 이미 나와 있는 논문을 찾아낸 것이었다.

그러니 "AI가 난제를 풀었다"는 말은 그 자체로는 근거가 못 된다.

이번엔 사람이 아니라 기계가 검증했다

오픈AI는 이번에 증명을 Lean 4 형식으로 만들어 깃허브에 공개했다. 아파치 2.0 라이선스다.

Lean 은 증명을 넣으면 컴파일되거나 안 되거나 둘 중 하나를 내놓는다. 설득의 여지가 없다.

그리고 저장소의 'sorry' 개수가 0이라고 한다. Lean에서 sorry 는 "이 단계는 일단 넘어간다"는 표시다. 0이면 빈 칸 없이 다 채웠다는 뜻이다.

모델을 믿을지 말지의 문제가, 컴파일러가 답하는 문제로 바뀐 셈이다.

무엇을 풀었나

대표 결과는 '비소픽군'의 명시적 구성이다. 미하일 그로모프가 1999년 소픽성 개념을 내놓은 뒤 27년간 열려 있던 질문이다.

나머지는 콘 강성 추측 반증, 고차원 구 채우기 밀도, 이진·구면 부호, 산술회로 복잡도, 양자 병렬 반복 같은 영역에 걸쳐 있다.

10건을 푸는 데 든 토큰 값은 GPT-5.6 Sol API 요금 기준 약 2,000달러로 계산됐다.

다만 밀레니엄 문제 — 클레이수학연구소가 2000년에 꼽은 7대 난제, 건당 상금 100만 달러 — 는 풀지 못했다.

그리고 논문으로 다듬은 건 사람이다. 오픈AI는 수학적 논증 자체는 모델에서 나왔다고 밝혔다.

AI 회사들이 죄다 연구실로 가고 있다

경향신문 보도를 보면 이건 오픈AI만의 움직임이 아니다.

  • 구글 딥마인드 — 5월에 에르되시 문제 353개 중 9개 해결. 지난달엔 알파폴드 전담팀을 해체하고 인력을 제미나이 등으로 재배치했다. 특화 모델에서 '범용 과학 에이전트'로 가는 방향이다
  • 앤트로픽 — 6월에 문헌 검색·코딩·데이터 분석·논문 초안을 한곳에서 처리하는 '클로드 사이언스'를 내놨다. 소외질환 신약 후보물질 발굴에도 나섰다
  • 오픈AI — 지난달 30일 학술기관 연구자 10만 명에게 첨단 모델을 무료로 주는 계획을 발표했다

이유는 여러 겹이다. 경향은 이렇게 짚는다.

회계 자동화 같은 건 일자리 감축 반발에 부딪히지만, 신약·신소재는 '인류 복지'라는 명분이 선다. 국가전략과 맞물려 정부 지원을 끌어내기도 좋다.

그리고 복잡한 연구 과정 자체가 모델 성능을 끌어올리는 시험장이 된다.

돈도 따라온다. 머크는 4월에 구글 클라우드와 최대 10억 달러 규모 계약을 맺었고, 글로벌마켓인사이트는 AI 신약 개발 시장이 올해 40억 달러에서 2035년 439억 달러로 커진다고 봤다.

한국은 인력이 스무 명 남짓이다

같은 보도에서 국내 상황도 나온다.

과기정통부가 지난달 AMD와 국가과학AI연구센터의 컴퓨팅 자원·인력 지원 협약을 맺었지만, 현재 인력은 20여 명 수준이다.

비슷한 역할을 하는 캐나다 밀라연구소는 교수 140명에 학생 연구자가 1,200여 명이다. 센터는 올 하반기에 인력을 더 뽑을 계획이라고 한다.

그래서 이게 나랑 무슨 상관이냐.

당장은 아니다. 비소픽군이 내 일에 쓰일 일은 없다.

다만 방식 하나는 가져갈 만하다. AI가 내놓은 결과를 기계가 확인할 수 있는 형태로 받는 것. 코드라면 테스트를 돌리고, 계산이라면 검산하고, 인용이라면 원문을 여는 식이다.

"맞는 것 같다"와 "컴파일된다"는 다르다. 이번 발표가 앞의 실패와 갈린 지점이 정확히 거기였다.

아스트라는 아직 안 나왔다. 출시일도 가격도 공개되지 않았다..

OpenAI says an unreleased model solved 10 long-open math problems.

The model is called Astra. The problems, in mathematics and theoretical computer science, had been open for a decade or more.

Claims like this are not new. Which is why the interesting part is elsewhere.

The same claim was made before, and rebutted

OpenAI's then vice president of science, Kevin Weil, once said GPT-5 had solved 10 previously unsolved Erdos problems.

Thomas Bloom, who maintains the Erdos problems database, called that "a dramatic misrepresentation." The model had not solved them; it had found papers that already contained the answers.

So "AI solved a hard problem" is not evidence on its own.

This time a machine did the checking

OpenAI formalized the proofs in Lean 4 and published them on GitHub under an Apache 2.0 license.

Lean returns a binary verdict: the proof compiles or it does not. There is nothing to argue about.

The repository reports a "sorry" count of zero. In Lean, sorry marks a step left unproven. Zero means no gaps.

Whether to trust the model became a question the compiler answers.

What it solved

The headline result is an explicit construction of a non-sofic group, open since Mikhail Gromov introduced soficity in 1999 — 27 years.

The rest spans a counterexample to Connes's rigidity conjecture, high-dimensional sphere packing density, binary and spherical codes, arithmetic circuit complexity, and quantum parallel repetition.

The token cost for all 10 solutions came to roughly $2,000 at GPT-5.6 Sol API rates.

It did not solve the Millennium Prize Problems — the seven the Clay Mathematics Institute named in 2000, each carrying a $1 million reward.

And humans turned the output into papers. OpenAI says the mathematical arguments themselves came from the model.

Every AI lab is moving into the research lab

Per the Kyunghyang Shinmun, this is not just OpenAI.

  • Google DeepMind — solved 9 of 353 Erdos problems in May. Last month it disbanded the dedicated AlphaFold team and moved staff to Gemini and other groups, shifting from single-purpose models to a general-purpose science agent
  • Anthropic — launched Claude Science in June, handling literature search, coding, data analysis and paper drafting in one place, and began work on drug candidates for neglected diseases
  • OpenAI — announced on July 30 that it would give 100,000 academic researchers free access to its frontier models

The reasons stack up, as the paper lays out.

Automating accounting invites a fight about jobs. Drug and materials discovery comes wrapped in public benefit, and it aligns with national strategy, which makes government support easier to attract.

Research also doubles as a testbed that pushes model capability.

The money follows. Merck signed a multiyear deal with Google Cloud worth up to $1 billion in April, and Global Market Insights projects the AI drug discovery market growing from $4 billion this year to $43.9 billion by 2035.

Korea's center has about twenty people

The same report covers the domestic picture.

Korea's science ministry signed an agreement with AMD last month for computing resources and staffing at the National Science AI Research Center. Its current headcount is around twenty.

Canada's Mila, which plays a comparable role, has 140 faculty and some 1,200 student researchers. The Korean center plans to hire more in the second half of this year.

So what does this mean for you?

Not much directly. Non-sofic groups will not come up in your work.

But the method transfers. Take AI output in a form a machine can check. Run the tests if it is code, redo the arithmetic if it is a calculation, open the source if it is a citation.

"Looks right" and "compiles" are different things. That is exactly where this announcement parted ways with the last one.

Astra is still unreleased. No date, no pricing..