AI 회사들이 헌책을 사들여 스캔한 뒤, 원본을 파쇄하고 있다고 한다.

영국 텔레그래프가 7월 30일 보도한 내용을 조선일보가 전했다. 나는 텔레그래프 원문을 직접 확인하지 못했으니, 조선일보가 인용한 범위에서만 쓴다.

먼저 왜 파쇄까지 하느냐가 이상하다. 스캔만 하고 책은 두면 되지 않나.

이유는 법에 있다는 게 보도의 설명이다.

디지털로 옮긴 뒤 원본을 완전히 없애면, 무단 복제가 아니라 종이를 디지털로 바꾼 '형식 변환'으로 볼 여지가 생긴다는 것. 저작권료를 내지 않고도 학습에 쓸 길이 열린다는 얘기다.

이건 확정된 판례가 아니라 그런 법적 해석이 있다는 보도다. 그대로 옮긴다.

실제로 드러난 사례가 있다.

앤트로픽은 7월 20일, 작가 3명이 저작권자들을 대표해 낸 집단소송에서 15억 달러(약 2조 2,000억 원)를 내기로 합의했다. 미국 저작권 소송 사상 최대 규모라고 한다.

그 과정에서 앤트로픽이 클로드 학습 데이터를 모으려고 중고 서점과 유통업체에서 종이책 수백만 권을 사들여 스캔한 뒤 재활용 업체로 보내 파기한 사실이 드러났다.

네덜란드·스페인·독일의 고서 판매상 수백 명이 올해 AI 업체들로부터 3,000종이 넘는 책을 한꺼번에 사겠다는 요청을 받았다.

텔레그래프는 매입된 책 상당수가 스캔을 마친 뒤 파쇄됐고, 그중에는 전 세계에 다섯 권밖에 남지 않은 희귀본도 있었다고 전했다.

다섯 권이 네 권이 됐다는 뜻이다.

왜 하필 종이책인가도 이유가 있다.

인터넷에 공개된 블로그·뉴스·위키피디아는 이미 거의 다 학습이 끝났다. 반면 종이책은 편집자와 출판사의 검증을 거친 글이라, 모델의 추론력과 문해력을 끌어올릴 자원으로 꼽힌다.

웹을 다 먹고 나니 도서관 쪽으로 눈이 간 셈이다.

학계 우려는 지식 독점 쪽이다. 보도에 인용된 한 업계 관계자는 빅테크가 희귀 서적을 쓸어 담으면 정작 그 책을 연구해야 할 학자들이 원본을 구하지 못하게 될 거라고 했다.

그래서 이게 나랑 무슨 상관이냐.

내가 쓰는 AI가 똑똑해진 데에는 누군가 사서 없앤 책이 섞여 있을 수 있다.

그게 곧바로 불법이라는 얘기는 아니다. 다만 스캔본은 회사 안에 남고 원본은 사라지는 구조라면, 그 지식에 접근할 수 있는 쪽이 한 곳으로 좁아진다.

AI가 아는 것과 내가 확인할 수 있는 것이 갈라지는 지점이다.

책이 사라진 자리에 남은 게 모델 가중치뿐이라면, 우리는 그걸 뭐라고 불러야 하나..

AI companies are buying up used books, scanning them, and shredding the originals.

The Telegraph reported this on July 30; the Chosun Ilbo carried it. I could not read the Telegraph piece directly, so I am working only from what Chosun quoted.

The shredding is the strange part. Why not scan the book and keep it?

The reason, per the report, is legal.

If the physical original is destroyed after digitization, the act can be framed not as unauthorized copying but as format-shifting from paper to digital — which opens a path to training on the text without paying royalties.

That is a reported legal reading, not settled case law. I am passing it along as such.

There is a documented case.

On July 20, Anthropic agreed to pay $1.5 billion to settle a class action brought by three authors on behalf of copyright holders — reportedly the largest settlement in the history of US copyright litigation.

In the process it emerged that Anthropic had bought millions of print books from used bookshops and distributors to gather training data for Claude, scanned them, and sent them to recyclers.

Hundreds of antiquarian booksellers in the Netherlands, Spain and Germany received requests this year from AI companies to buy more than 3,000 titles at once.

The Telegraph reported that many of the purchased books were shredded after scanning, including copies of works with only five known to exist worldwide.

Five became four.

There is a reason it is paper books specifically.

Blogs, news and Wikipedia on the open web have largely been consumed already. Print books carry text that passed through editors and publishers, which makes them a prized resource for raising a model's reasoning and reading ability.

Having finished the web, they turned toward the library.

The academic concern is concentration. An industry figure quoted in the report said that if big tech sweeps up rare books, the scholars who need to study them will no longer be able to find originals.

So what does this mean for you?

Some of what makes your AI capable may include books someone bought and destroyed.

That is not automatically illegal. But if the scan stays inside a company while the original disappears, access to that knowledge narrows to one place.

That is where what the AI knows and what you can verify start to diverge.

If all that remains where a book stood is a set of model weights, what should we call that..

Sources · The Chosun Ilbo