On April 30th, Marçal Font received an order that caught his attention at Librería Fénix, a landmark for the trade of antique books in the Catalan city of Badalona.
"I've been a secondhand bookseller for 20 years, so I've had millions of strange purchases, which means I'm not alarmed anymore," he tells BBC Mundo. He is also a poet and professor of Comparative Literature at the University of Barcelona.
The order came from a company located in Canada, but this wasn't the first time the bookstore had shipped Catalan books abroad. "We send books as far as Japan, it's not unusual; but the (purchase) pattern was strange."
When a new series of orders arrived at the beginning of May, Font decided to investigate for fear that it was a fraud.
"It's not common to look for who buys from us, but in that case we started looking, mainly because the orders came in a window of one hour or 40 minutes. We had three or four emails asking for a random number of books. And the associated shipping costs were outrageous."
He then consulted with his friend Xavier Vinaixa, with whom he shares not only a love for books and humanism, but also a passion for artificial intelligence.
"He called me and said, 'Hey, I'm getting orders here with strange patterns, from the same client, with international shipping and paying three times the shipping costs.' And his first thought was that it might be a scam," Vinaixa recalls in an interview with BBC Mundo.
"And that's when we started pulling on the thread and found other people, with the same concerns as Marçal, in Germany, in New Zealand... And then we saw it very clearly," adds this artificial intelligence researcher and technical director of the technology company Sorensen.ai.
If one performs an internet search with the terms "old books" and "artificial intelligence" one quickly finds those "concerns" to which Vinaixa refers.
In an article in the British newspaper The Telegraph entitled "Cultural barbarism: how AI companies are destroying the world's books", a bookseller from Haarlem, Netherlands, recounts the surprise he felt upon receiving a massive order of books:
"In the antiquarian book trade, it's very rare for people to want to buy more than one book. So, if someone comes in and says 'I want a couple of hundred of your books,' it's very strange," said antiquarian book dealer Pieter de Vries.
The Guardian quoted a bookseller in the Australian city of Victoria who received an order from the same company that bought from Marçal Font: "They placed the order, paid in advance and didn't object to the shipping costs," explained Delfina Manor.
Both articles, as well as those that told the story of the Catalan bookseller, refer to the journalistic investigation that uncovered what is known as the "Panama Project".
In September 2025, the artificial intelligence (AI) company Anthropic agreed to pay US$1.5 billion to settle a class-action lawsuit filed by authors who alleged that the company had used their works without permission to train its AI models.
Four months later, the Washington Post exclusively published a document that was part of the legal proceedings, which mentioned the existence of a plan to digitize all the works written by humanity.
"The Panama Project is our effort to destructively scan all the books in the world," the leaked text stated, later specifying: "We don't want it to be known that we are working on this."
Maximiliano Firtman, an Argentine professor and author of books on programming and artificial intelligence, was one of the authors who sued Anthropic in that class action lawsuit.
According to what he told BBC Mundo, the judge ruled that if the company bought a copy of the book, it could use it to train the AI.
"And that's when the testimonies began to emerge from booksellers around the world who received orders for books that no one touched, and from some companies that are dedicated to scanning these books and giving the scanned text to artificial intelligence companies. That's how the issue of the destruction of books arises."
Rodrigo Álvarez, a journalist specializing in technology at the Argentine news outlet Todo Noticias (TN), explains that there are two methods for digitizing books, one of which does not involve the destruction of the copy, since it uses flatbed scanners or V-shaped devices that scan the pages without disassembling the binding.
"The 'destructive' procedure is basically a process that begins with cutting the spine so that the sheets are loose; the separated pages are then fed into high-speed industrial scanners that scan them much faster, automatically," he tells BBC Mundo.
Once the pages are digitized, the book can no longer be reassembled and its remains are usually sent for recycling.
"Companies choose this second alternative for reasons of scale: it allows for the digitization of millions of copies more quickly and at a lower cost. In the documented case of Anthropic, this procedure was used for the Panama Project," Álvarez concludes.
Digitizing books allows us to preserve texts written throughout human history, but the main objective seems to be more about training artificial intelligence than about a spirit of conservation.
"When we talk about AI, it is usually understood that we are talking about an LLM ( Large Language Model ), a large language model," explains Xavier Vinaixa.
LLMs, such as ChatGPT, are artificial intelligence programs that learn to understand, summarize, translate, and generate text or other content by predicting the next word or fragment.
"We all use predictive text on our phones. You say 'hello, I'm arriving...' and it might suggest the word 'late'. How does that predictive text work? Based on statistics from previous texts you've written. That same thing, but raised to the millionth power, is an LLM," explains Maximiliano Firtman.
So, to train these large language models in the intricate art of finding connections between words, they were first "fed" with all the information available on the web - Wikipedia, blogs, forums - and when this began to run out, they were trained with newspaper articles and even subtitles from YouTube videos.
"When we talk about AI, it is usually understood that we are talking about an LLM ( Large Language Model ), a large language model," explains Xavier Vinaixa.
LLMs, such as ChatGPT, are artificial intelligence programs that learn to understand, summarize, translate, and generate text or other content by predicting the next word or fragment.
"We all use predictive text on our phones. You say 'hello, I'm arriving...' and it might suggest the word 'late'. How does that predictive text work? Based on statistics from previous texts you've written. That same thing, but raised to the millionth power, is an LLM," explains Maximiliano Firtman.
So, to train these large language models in the intricate art of finding connections between words, they were first "fed" with all the information available on the web - Wikipedia, blogs, forums - and when this began to run out, they were trained with newspaper articles and even subtitles from YouTube videos.
The destruction of the books scanned to train the AI has generated questions and doubts from various forums.
Miguel Ángel Ortega, president of the Professional Association of Antique Books and Collectibles of Spain (UNILIBER), tells BBC Mundo that a distinction should be made between copies from mass print runs and limited print runs.
"As merchants who preserve heritage and carry out conservation work, we believe that the final use of a work, of a very limited, scarce or unique edition, logically, should not be destruction."
Rodrigo Álvarez, the Technology journalist for TN, agrees, stating that the biggest risk "is that indiscriminate purchases will destroy rare or out-of-print copies without first checking how many of those copies remain in stock at distributors and bookstores, or even in the possession of individuals."
Marçal Font, on the other hand, is not so worried about the destruction of books: "I'm not at all romantic in that sense."
"Just as veterinarians who study to save animals end up being the ones who sacrifice the most animals; my point of view is that there is no one who destroys more books than a secondhand bookseller, because we buy entire libraries, we save everything that can be saved, but there is much that cannot be saved."
To avoid digitization that destroys books and leaves out key information, both Font and Vinaixa propose avoiding intermediaries.
"The solution we propose is to scan our heritage in our own language, store it, and license the data; if you want the data, I'll give it to you in the best way possible, and I'll keep the book, because for people who do research, it's also important how the book has its cover, how it has been sewn, not just the content," says Vinaixa.
The technical director of the technology company Sorensen.ai adds that the demand for data will continue until there is no more data to feed artificial intelligence.
"We thought AI would become too small when there was no more processing power. Well, no, the limit is data; AI has become too small because there is no more data."
According to Firtman, we still don't know what will happen when human-generated information runs out:
"The next thing we have is synthetic data, texts written by the AI itself. The biggest risk in a few decades is that when we ask everything to an AI, without going to the sources, that AI will start to degrade because it is being trained with the same information that the AI itself generated."
Journalist Rodrigo Álvarez illustrates this degradation with the image of a snake eating its own tail:
"Companies would feed their models with content generated by other artificial intelligences, and in that case, errors could multiply, making the language increasingly basic, like a summary of a summary of a summary, increasingly simple, a synthesis of a synthesis that moves away from the original information."
Wednesday, August 12, 2026
What is the "Panama Project" and how did it reveal the destruction of thousands of ancient books to train AI?
On April 30th, Marçal Font received an order that caught his attention at Librería Fénix, a landmark for the trade of antique books in the Catalan city of Badalona.
"I've been a secondhand bookseller for 20 years, so I've had millions of strange purchases, which means I'm not alarmed anymore," he tells BBC Mundo. He is also a poet and professor of Comparative Literature at the University of Barcelona.
The order came from a company located in Canada, but this wasn't the first time the bookstore had shipped Catalan books abroad. "We send books as far as Japan, it's not unusual; but the (purchase) pattern was strange."
When a new series of orders arrived at the beginning of May, Font decided to investigate for fear that it was a fraud.
"It's not common to look for who buys from us, but in that case we started looking, mainly because the orders came in a window of one hour or 40 minutes. We had three or four emails asking for a random number of books. And the associated shipping costs were outrageous."
He then consulted with his friend Xavier Vinaixa, with whom he shares not only a love for books and humanism, but also a passion for artificial intelligence.
"He called me and said, 'Hey, I'm getting orders here with strange patterns, from the same client, with international shipping and paying three times the shipping costs.' And his first thought was that it might be a scam," Vinaixa recalls in an interview with BBC Mundo.
"And that's when we started pulling on the thread and found other people, with the same concerns as Marçal, in Germany, in New Zealand... And then we saw it very clearly," adds this artificial intelligence researcher and technical director of the technology company Sorensen.ai.
If one performs an internet search with the terms "old books" and "artificial intelligence" one quickly finds those "concerns" to which Vinaixa refers.
In an article in the British newspaper The Telegraph entitled "Cultural barbarism: how AI companies are destroying the world's books", a bookseller from Haarlem, Netherlands, recounts the surprise he felt upon receiving a massive order of books:
"In the antiquarian book trade, it's very rare for people to want to buy more than one book. So, if someone comes in and says 'I want a couple of hundred of your books,' it's very strange," said antiquarian book dealer Pieter de Vries.
The Guardian quoted a bookseller in the Australian city of Victoria who received an order from the same company that bought from Marçal Font: "They placed the order, paid in advance and didn't object to the shipping costs," explained Delfina Manor.
Both articles, as well as those that told the story of the Catalan bookseller, refer to the journalistic investigation that uncovered what is known as the "Panama Project".
In September 2025, the artificial intelligence (AI) company Anthropic agreed to pay US$1.5 billion to settle a class-action lawsuit filed by authors who alleged that the company had used their works without permission to train its AI models.
Four months later, the Washington Post exclusively published a document that was part of the legal proceedings, which mentioned the existence of a plan to digitize all the works written by humanity.
"The Panama Project is our effort to destructively scan all the books in the world," the leaked text stated, later specifying: "We don't want it to be known that we are working on this."
Maximiliano Firtman, an Argentine professor and author of books on programming and artificial intelligence, was one of the authors who sued Anthropic in that class action lawsuit.
According to what he told BBC Mundo, the judge ruled that if the company bought a copy of the book, it could use it to train the AI.
"And that's when the testimonies began to emerge from booksellers around the world who received orders for books that no one touched, and from some companies that are dedicated to scanning these books and giving the scanned text to artificial intelligence companies. That's how the issue of the destruction of books arises."
Rodrigo Álvarez, a journalist specializing in technology at the Argentine news outlet Todo Noticias (TN), explains that there are two methods for digitizing books, one of which does not involve the destruction of the copy, since it uses flatbed scanners or V-shaped devices that scan the pages without disassembling the binding.
"The 'destructive' procedure is basically a process that begins with cutting the spine so that the sheets are loose; the separated pages are then fed into high-speed industrial scanners that scan them much faster, automatically," he tells BBC Mundo.
Once the pages are digitized, the book can no longer be reassembled and its remains are usually sent for recycling.
"Companies choose this second alternative for reasons of scale: it allows for the digitization of millions of copies more quickly and at a lower cost. In the documented case of Anthropic, this procedure was used for the Panama Project," Álvarez concludes.
Digitizing books allows us to preserve texts written throughout human history, but the main objective seems to be more about training artificial intelligence than about a spirit of conservation.
"When we talk about AI, it is usually understood that we are talking about an LLM ( Large Language Model ), a large language model," explains Xavier Vinaixa.
LLMs, such as ChatGPT, are artificial intelligence programs that learn to understand, summarize, translate, and generate text or other content by predicting the next word or fragment.
"We all use predictive text on our phones. You say 'hello, I'm arriving...' and it might suggest the word 'late'. How does that predictive text work? Based on statistics from previous texts you've written. That same thing, but raised to the millionth power, is an LLM," explains Maximiliano Firtman.
So, to train these large language models in the intricate art of finding connections between words, they were first "fed" with all the information available on the web - Wikipedia, blogs, forums - and when this began to run out, they were trained with newspaper articles and even subtitles from YouTube videos.
"When we talk about AI, it is usually understood that we are talking about an LLM ( Large Language Model ), a large language model," explains Xavier Vinaixa.
LLMs, such as ChatGPT, are artificial intelligence programs that learn to understand, summarize, translate, and generate text or other content by predicting the next word or fragment.
"We all use predictive text on our phones. You say 'hello, I'm arriving...' and it might suggest the word 'late'. How does that predictive text work? Based on statistics from previous texts you've written. That same thing, but raised to the millionth power, is an LLM," explains Maximiliano Firtman.
So, to train these large language models in the intricate art of finding connections between words, they were first "fed" with all the information available on the web - Wikipedia, blogs, forums - and when this began to run out, they were trained with newspaper articles and even subtitles from YouTube videos.
The destruction of the books scanned to train the AI has generated questions and doubts from various forums.
Miguel Ángel Ortega, president of the Professional Association of Antique Books and Collectibles of Spain (UNILIBER), tells BBC Mundo that a distinction should be made between copies from mass print runs and limited print runs.
"As merchants who preserve heritage and carry out conservation work, we believe that the final use of a work, of a very limited, scarce or unique edition, logically, should not be destruction."
Rodrigo Álvarez, the Technology journalist for TN, agrees, stating that the biggest risk "is that indiscriminate purchases will destroy rare or out-of-print copies without first checking how many of those copies remain in stock at distributors and bookstores, or even in the possession of individuals."
Marçal Font, on the other hand, is not so worried about the destruction of books: "I'm not at all romantic in that sense."
"Just as veterinarians who study to save animals end up being the ones who sacrifice the most animals; my point of view is that there is no one who destroys more books than a secondhand bookseller, because we buy entire libraries, we save everything that can be saved, but there is much that cannot be saved."
To avoid digitization that destroys books and leaves out key information, both Font and Vinaixa propose avoiding intermediaries.
"The solution we propose is to scan our heritage in our own language, store it, and license the data; if you want the data, I'll give it to you in the best way possible, and I'll keep the book, because for people who do research, it's also important how the book has its cover, how it has been sewn, not just the content," says Vinaixa.
The technical director of the technology company Sorensen.ai adds that the demand for data will continue until there is no more data to feed artificial intelligence.
"We thought AI would become too small when there was no more processing power. Well, no, the limit is data; AI has become too small because there is no more data."
According to Firtman, we still don't know what will happen when human-generated information runs out:
"The next thing we have is synthetic data, texts written by the AI itself. The biggest risk in a few decades is that when we ask everything to an AI, without going to the sources, that AI will start to degrade because it is being trained with the same information that the AI itself generated."
Journalist Rodrigo Álvarez illustrates this degradation with the image of a snake eating its own tail:
"Companies would feed their models with content generated by other artificial intelligences, and in that case, errors could multiply, making the language increasingly basic, like a summary of a summary of a summary, increasingly simple, a synthesis of a synthesis that moves away from the original information."
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment