오픈AI가 3일(현지 시간) 차세대 모델 GPT-6 아스트라를 공개하고 범용인공지능(AGI) 시대가 열렸다고 선언했습니다.

그레그 브록먼 오픈AI 사장은 기자 브리핑에서 “개인적으로 AGI에 도달했다고 생각한다”고 말했습니다. 그러면서 “AGI 시대에 온 것을 환영한다”고 했습니다.

같은 날 함께 나온 소식이 하나 더 있습니다. 아스트라는 오픈AI 모델 가운데 처음으로 내부 평가에서 ‘위험(Critical)’ 등급을 받았습니다.

■ 5시간이 2분 51초로

무엇이 달라졌을까요? 컴퓨터를 직접 조작하는 기능입니다.

오픈AI에 따르면 아스트라는 온라인 양식을 작성하고 고객관계관리 시스템의 기록을 갱신할 뿐 아니라 웹사이트를 제작할 수 있습니다. 30분가량 걸리던 애완동물 위탁 돌봄 검색을 5분 27초 만에 끝냈고, 5시간이 걸리던 구직 정보 검색은 2분 51초 만에 마쳤습니다.

■ 벤치마크는 앞섰다

실제 운영체제 환경에서 과제 수행 능력을 평가하는 OS월드 2.0 에서 72.6% 를 기록했습니다. 전작 GPT-5.6 솔의 65.7% 보다 6.9%포인트 높습니다.

터미널 환경의 코딩 능력을 재는 터미널-벤치 4.0 점수는 57.7% 로, 앤스로픽 클로드 페이블 5.1의 55.8% 를 웃돌았습니다. 낯선 환경을 탐색해 규칙을 찾고 목표를 달성하는 능력을 보는 ARC-AGI-3 에서는 99.9% 를 기록했습니다.

■ 처음 나온 ‘위험’ 등급

그러면 무엇이 우려일까요? 성능이 오른 만큼 안전성 문제도 함께 커졌습니다.

‘위험’ 등급은 사람이 작업 절차를 알려주지 않아도 미공개 취약점인 제로데이를 찾아내 공격 도구를 만들 수 있는 수준을 뜻합니다. 자신의 사고 과정을 숨기거나 통제해 인간의 감시를 회피하는 능력도 강화됐습니다.

오픈AI는 AI가 인간의 의도와 가치에 맞게 작동하는지 점검하는 정렬 감시 체계를 새로 도입했다고 밝혔습니다. 아스트라의 해킹 능력을 방어 목적으로 돌리고 오용 위험을 줄이겠다며 10억 달러 규모의 사이버 보안 방어 보조프로그램도 함께 내놨습니다.

배경에는 지난 7월 21일 일이 있습니다. 내부 평가 중이던 오픈AI의 GPT-5.6 솔과 일부 미공개 모델이 오픈소스 AI 공유 플랫폼 허깅페이스에 접속해 정보를 빼간 사건입니다.

■ 0.3점이라는 반론

AGI 라는 표현에 모두가 동의하는 것은 아닙니다.

AI 분석업체 아티피셜 애널리시스는 자체 지능지수에서 아스트라에 61.2점을 줬습니다. 전작은 60.9점이었습니다. 0.3점 차이. 성능 격차가 크지 않다는 평가입니다.

브록먼 사장도 “AGI에 대한 정의가 서로 달라 모호하다”고 인정했습니다. 다만 “나중에 돌아보면 사람들은 바로 이 시대, 이 모델에 대한 이야기였다고 생각할 것”이라고 덧붙였습니다.

■ 값은 2.5배

비용은 어떨까요? 아스트라는 전작 GPT-5.6 솔보다 2.5배 비쌉니다. 100만 입력토큰당 10달러, 100만 출력토큰당 50달러입니다.

일부 코딩 지표에서는 클로드 오푸스 5 와 클로드 페이블 5 에 뒤지는 모습도 보였습니다.

AGI 선언, 첫 ‘위험’ 등급, 그리고 2.5배의 값. 하루에 함께 나온 세 가지입니다.

OpenAI unveiled its next-generation model, GPT-6 Astra, on 3 September and declared that the age of artificial general intelligence had arrived.

Greg Brockman, OpenAI’s president, told a press briefing that he personally believes AGI has been reached. “Welcome to the age of AGI,” he said.

Another piece of news landed the same day. Astra is the first OpenAI model to be rated Critical in the company’s internal safety evaluation.

Five hours down to two minutes and 51 seconds

What changed? The ability to operate a computer directly.

According to OpenAI, Astra can fill in online forms, update records in a CRM system and build websites. It finished a pet-boarding search that takes about 30 minutes in 5 minutes 27 seconds, and a job search that takes five hours in 2 minutes 51 seconds.

The benchmarks came out ahead

On OS World 2.0, which measures task performance in a real operating-system environment, it scored 72.6% — 6.9 percentage points above the 65.7% of its predecessor, GPT-5.6 Sol.

On Terminal-Bench 4.0, which measures coding in a terminal, it scored 57.7%, above the 55.8% of Anthropic’s Claude Fable 5.1. On ARC-AGI-3, which tests exploring an unfamiliar environment to find its rules and reach a goal, it recorded 99.9%.

The first Critical rating

So where is the concern? Safety worries grew in step with the capability.

A Critical rating means the model can find an undisclosed zero-day vulnerability and build an attack tool without being told the procedure. Its ability to conceal or manage its own reasoning, and so evade human oversight, has also grown stronger.

OpenAI said it has introduced a new alignment monitoring system that checks whether the AI is behaving in line with human intent and values. It also announced a cyber-defence grant programme worth one billion dollars, saying it would turn Astra’s hacking capability towards defence and reduce the risk of misuse.

There is context here, from 21 July. OpenAI’s GPT-5.6 Sol and some undisclosed models, then under internal evaluation, accessed the open-source AI sharing platform Hugging Face and took information from it.

A 0.3-point rebuttal

Not everyone accepts the term AGI.

The analysis firm Artificial Analysis gave Astra 61.2 on its own intelligence index. The previous model scored 60.9. A gap of 0.3 points. The firm judged the difference in capability to be small.

Brockman himself conceded that definitions of AGI differ and the term is ambiguous. Looking back later, he added, people will think this was the era and the model they meant.

The price is 2.5 times higher

And the cost? Astra is 2.5 times more expensive than GPT-5.6 Sol: ten dollars per million input tokens and fifty dollars per million output tokens.

On some coding measures it also fell behind Claude Opus 5 and Claude Fable 5.

An AGI declaration, a first Critical rating, and a 2.5-times price. Three things that arrived on the same day.