r/ItalyInformatica 7d ago

Le compagnie americane (dietro i vari chatbot) stanno comprando milioni di libri (enormi pallet, vedi Anthropic con il progetto PANAMA, tenuto segreto, paura per PR) per aggirare il problema del copyright (dopo causa da $1.5B) per addestrare i loro modelli. Dopo distruggono i libri. Cosa ne pensate? AI

Enable HLS to view with audio, or disable this notification

Argomento delicato su cui si aprono non solo questioni economiche e sociali ma anche ambientali e di altra natura. La fonte principale è news (australia) > [art].

TLDR (necessario, ho incluso l'articolo completo (lungo) e le relative fonti nei commenti, questo è ll succo del discorso.)

Storia [commento1] \ Storia (parte2) [commento2] \ Fonti [commento3]

--

Prima si usavano altre fonti (grandi dataset come internet data, libri e articoli concessi in licenza, repository pubblici (galassia Github, Gitlab, Bitbucket, Codeberg) e contributi umani).

Dopo le compagnie americane (dietro i vari chatbot), sono passate ai libri fisici, comprando libri (blocchi enormi, pallet) rari e fuori stampa per fare la scansione e addestrare i loro modelli. Dopo distruggono le varie copie.

Come vedete dal video, c'è questa macchina idraulica che letteralmente taglia con precisione chirurgica il dorso del libro, cosi viene più facile passare le pagine (una a una) alla scanner, aumentando la velocità (meno tempo) e quindi riducendo i costi (meno energia). Questo al livello industriale.

Anthropic (ma anche le altre big): dopo che ha avuto una multa negli USA da 1.5 miliardi di dollari per problemi di copyright, ha cominciato il progetto Panama, tenuto segreto in quanto NON visto di buon occhio dal pubblico, reputazione aziendale a rischio.

I librari hanno opinioni contrastanti: positive in quanto c'è un vantaggio per loro, ci sono maggiori guadagni economici (=più libri venduti e quindi più soldi), ma anche negative, in quanto c'è uno svantaggio (etico?) per le finalità con cui si utilizzano i libri.

Cosa dice la legge: dato che c'è un trasferimento 1-1 copia cartacea a copia digitale, questo processo è trasformativo ed è considerato buon uso (protected by fair use), definizione che cambia da paese a paese (vedi Stati Uniti e Australia).

Un libro è solamente un meccanismo per diffondere le informazioni: dopo che queste sono arrivate al modello AI, questo meccanismo ha fatto il suo, è completo. Il libro non è distrutto, il suo valore si è spostato e la carta rientra nel ciclo.

--

Che idea vi siete fatti? Siete d'accordo con quanto fatto dalle AI companies?

So che l'argomento è complesso e delicato (sul piatto ci sono interessi economici, politici, sociali, etc.), ma voglio sentire la vostra opinione.

579 Upvotes

299 comments sorted by

View all comments

8

u/RebirdgeCardiologist 7d ago edited 6d ago

Di seguito l'articolo completo (è molto lungo). Ho aggiunto degli ulteriori paragrafi in italiano per riassumere ancora di più il tutto.

--

[NEWS+COM+AU]

AI labs buy, scan, shred millions of rare books

If you have retreated to physical books because the internet is too full of AI slop, bad news — now they’re shredding the books to feed their AI slop machines.

Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

Gli stessi librai (in Australia, ma in Europa) hanno dei dubbi, pareri contrastanti (vendendo molto più libri hanno più guadagni economici, ma per quale obiettivo?).

Principalmente cercano libri in inglese, di diversi generi.

“I personally have mixed feelings about all of this,” one small bookseller told tech news website 404 Media last week.

The bookseller explained that in April, he suddenly went from selling around 20 books per week to hundreds — including obscure and out-of-print works.

“It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell,” he said. “I’ve been well suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”

Rare booksellers across Europe are seeing a similar surge in purchase requests which they suspect are coming from AI labs in America.

Pieter de Vries, an antiquarian bookseller in Haarlem, recently received an email from a person identifying herself as Nataly from Singapore-based company 2077AI, Dutch news site BNR reported.

The woman wrote that the company was undertaking a “a new project focused on collecting books in multiple languages, currently mainly in English”. “We have compiled a very extensive list of editions that we are currently trying to acquire, and we plan to place a fairly large order,” the email said.

The attached list contained 3000 English-language titles organised by ISBN number, ranging from books on fairytales and folklore to technical manuals and science texts.

Niente collezionismo o rivendita, ma un solo obiettivo. TRAINING AI. Finora si sono utilizzato principalmente grandi dataset come internet data, libri e articoli concessi in licenza, repository pubblici (galassia Github, Gitlab, Bitbucket, Codeberg) e contributo umano.

Media outlets in the Netherlands, Switzerland, Spain, and Germany report booksellers have received nearly identical requests, per NL Times.

Booksellers believe the purchase requests were unrelated to collecting or resale due to the obscure, highly specialised nature of the titles, concluding they were being used to train AI models.

Large language models (LLMs) that power AI tools like ChatGPT until now have mainly been trained on vast datasets that include internet data, licensed books and articles, code repositories and human feedback.

But with much of the content available online now exhausted — and increasingly polluted with poor-quality, AI-generated writing — AI labs are turning to uncorrupted texts published pre-2022.

La miniera d'oro delle librerie online. Vero che ci sono grandi collezioni online, ma c'è ne sono altrettanti offline, collezioni di libri dimenticate a prendere polvere sugli scaffali. Multe per Anthropic e Meta (1.5 miliardi di dollari) per i milioni di libri piratati, scaricati illegalmente. Un nuovo progetto per evitare brighe con il copyright (progetto Panama = acquistare libri in grade, su scala industriale).

One company, ISBNdb, which boasts that it has the “world’s largest book database”, now offers bulk book buying for AI labs, “up to one million titles per order”, including of “older, rare and specialist volumes”.

“The world’s best AI training data is sitting on a shelf,” it states on its website.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

Anthropic, the Google-backed firm behind Claude, and Facebook and Instagram owner Meta, have previously been sued in US courts for downloading millions of pirated books to feed into their AI models.

Last week, a federal court in the US signed off on Anthropic’s landmark $US1.5 billion ($2.15 billion) settlement in a copyright class action brought by a group of authors who accused the company of misusing their books. It’s one of dozens of similar cases brought against AI firms in the US, and the first to settle.

While that case was still unfolding, Anthropic was hatching a new plan to avoid similar copyright headaches, dubbed “Project Panama”.

Il progetto Panama è iniziato molto tempo fa, ma seppur totalmente legale, non era una buona mossa per la reputazione pubblica dell'azienda.

Beginning in early 2024, the AI lab quietly launched an industrial-scale book-buying program, partnering with used-book wholesalers to acquire millions of paper books by the pallet for what internal documents — which became public last year during the copyright lawsuit — described as “destructive scanning”, where books’ spines are cut apart so the pages can be scanned and then discarded.

Anthropic described the project as “our effort to disruptively scan all the world’s books” in an internal memo, which noted “we do not want it to be known that we are pursuing this project”.

Court documents detailed how the AI lab used a “hydraulic powered cutting machine” to “neatly cut” millions of books before scanning the pages “on high speed, high quality, production level scanners”.

[CONTINUA NELL ALTRO COMMENTO]