<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Copyright | Clément Godbarge</title><link>https://clementgodbarge.com/tag/copyright/</link><atom:link href="https://clementgodbarge.com/tag/copyright/index.xml" rel="self" type="application/rss+xml"/><description>Copyright</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 06 Sep 2026 00:00:00 +0000</lastBuildDate><image><url>https://clementgodbarge.com/media/icon_hu70ea0d613120738a4eb6fafa2b227dfa_50891_512x512_fill_lanczos_center_3.png</url><title>Copyright</title><link>https://clementgodbarge.com/tag/copyright/</link></image><item><title>The Fence, Not the Fire</title><link>https://clementgodbarge.com/post/fence-not-fire/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://clementgodbarge.com/post/fence-not-fire/</guid><description>&lt;p>In a warehouse in North Las Vegas, Amazon runs a machine that cuts the spines off second-hand books. The loose pages go through a scanner, and then the paper goes off to be pulped. The &lt;a href="https://www.dailymail.com/news/article-16088549/The-modern-book-burners-tech-giants-buying-thousands-second-hand-books-destroy-FRED-KELLY.html" target="_blank" rel="noopener">Daily Mail&lt;/a> compared this activity to modern book burning, sparing no clichés in the process: allusions to Savonarola, stock photos of seemingly rare books and, of course, the Berlin pyre. As expected, the broadsheets were cooler, but left the same comparison hanging in the air. For all their indignation, these articles let the AI industry off remarkably lightly. Who gets to control the knowledge being scanned? That question will have to wait while we dispose of the book-burning nonsense.&lt;/p>
&lt;h2 id="a-scanner-is-not-a-bonfire">A scanner is not a bonfire&lt;/h2>
&lt;p>Destroying a physical copy does not, by itself, make destructive scanning a modern equivalent of Savonarola’s bonfires or Nazi book burning. Those acts drew their meaning from what was condemned and why. That poor monk Savonarola, dragged into every book-burning controversy, called on people to burn some of their possessions as a public act of penance. The Nazi bonfires staged a different kind of repudiation: burning the books publicly declared that their authors had no place in German cultural life. In both cases, destruction was performative. Destructive scanning, a widely used method in digitisation, serves a different purpose: the copy is sacrificed to retain its contents. The book is destroyed; the text is dematerialised. Anyone reaching for &lt;em>Fahrenheit 451&lt;/em> might recall that Bradbury’s firemen were not saving the text for later.&lt;/p>
&lt;p>&lt;a href="https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/" target="_blank" rel="noopener">404 Media&lt;/a> added to the alarm by putting ‘rare books’ in its headline, a description other outlets duly repeated. Yet a bookseller cited in its own investigation said these bulk orders &lt;a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/secret-tracking-device-placed-in-rare-book-ends-up-in-amazon-processing-facility-destroying-books-to-train-ai-models-is-all-the-vegas-warehouse-does" target="_blank" rel="noopener">excluded books without ISBNs&lt;/a>, pointing largely to editions published since the introduction of ISBNs around 1970. This is the stuff of car boot sales and charity shops, often one clear-out away from the tip. Before announcing a cultural catastrophe, one might check the library catalogue. These are precisely the publications that legal-deposit libraries exist to preserve. Admittedly, an ISBN does not guarantee a surviving deposit copy, but neither does scarcity on the second-hand market establish a text’s disappearance.&lt;/p>
&lt;p>Pulping destroys the physical copy, including any inscriptions or marginalia. Copies with historically significant annotations or provenance deserve to be identified and preserved, ideally before they reach the second-hand market. We might still worry about the dwindling supply of physical copies of out-of-print titles. But in that case, we should also ask why their rights holders neither reprint them nor authorise access to existing digital facsimiles. Scholars know the frustration: Google Books finds the passage you need, then offers a snippet of a book you can neither buy new nor read online. There are losses here worth debating. They do not turn a scanner into a Nazi bonfire.&lt;/p>
&lt;p>The book trade, meanwhile, has long pulped unsold books as a matter of routine. In France alone, an estimated &lt;a href="https://actualitte.com/article/120833/edition/livres-invendus-un-taux-de-retour-de-22-et-25-000-tonnes-pilonnees" target="_blank" rel="noopener">25,000 tonnes of books were pulped in 2023&lt;/a>, roughly 60 per cent of the tonnage returned unsold to distributors. This is called managing stock. Surplus copies and scarce second-hand titles raise different questions of access, but destroying paper does not in itself amount to censorship or tyranny.&lt;/p>
&lt;h2 id="more-books-better-ai">More books, better AI?&lt;/h2>
&lt;p>That AI companies are buying and scanning books is, in fact, good news. A technology routinely criticised for its shallow grasp of human culture is acquiring more of it, and one might have expected scholars to welcome the effort.&lt;/p>
&lt;p>The web is a poor substitute for the accumulated record of a language. Much of a language’s recorded life, from novels and memoirs to popular songs and transcribed conversations, has never made it from print to the web. Many of these books have outlived their publishers. Bringing them into training data matters most where little else is available. A few more books in English may bring some marginal improvements. In a poorly represented language, those books may make the difference between a usable system and a useless one. Whether these tools belong in our public services or everyday lives is a question for public debate. But wherever we choose to use them, we should be able to do so in our own language.&lt;/p>
&lt;p>The orders reported in Spain and the Netherlands make the linguistic stakes concrete. In May, &lt;a href="https://www.eldiario.es/catalunya/misteriosa-empresa-compra-libros-viejos-entrenar-ia-destruye-expolio-literario_1_13235821.html" target="_blank" rel="noopener">elDiario.es&lt;/a> described a Badalona bookseller receiving repeated orders, chiefly for Catalan non-fiction: history books, technical manuals and old conference proceedings. By August, Dutch reporting described a request to De Slegte for &lt;a href="https://www.bnr.nl/nieuws/tech-innovatie/10607371/de-slegte-krijgt-mysterieuze-megabestelling-zoom-books-wil-800-000-nederlandse-boeken" target="_blank" rel="noopener">800,000 Dutch-language books&lt;/a>. The language of the books deserves at least as much attention as the fate of their bindings. Perhaps that is too much to ask of the Daily Mail. Even the more sober papers have had little to say about the languages involved.&lt;/p>
&lt;p>Licensing agreements with large publishers can also supply valuable material, but their digital catalogues cover only a fraction of what has appeared in print. Books never digitised, or whose rights the publisher no longer holds, fall outside the offer. Even a generous licensing deal cannot stand in for the history of publishing, still less for the history of a culture. Buying second-hand copies is a way of reaching past those commercial boundaries.&lt;/p>
&lt;h2 id="theft-is-not-an-argument">‘Theft’ is not an argument&lt;/h2>
&lt;p>Calling it ‘theft’ does not help either. In June 2025, two US district courts held that using books to train language models was fair use in the cases before them. In &lt;a href="https://copyrightalliance.org/wp-content/uploads/2025/06/Bartz-v.-Anthropic-Order.pdf" target="_blank" rel="noopener">Bartz v. Anthropic&lt;/a>, the court stressed the transformative character of training: the books were used to build a system capable of producing new expression. Models can memorise passages, but using a book for training does not make the whole text retrievable on demand. Nor does the absence of verbatim reproduction dispose of every objection. In &lt;a href="https://law.justia.com/cases/federal/district-courts/california/candce/3:2023cv03417/415175/598/" target="_blank" rel="noopener">Kadrey v. Meta&lt;/a>, the judge warned that AI-generated works could flood the market and that evidence of such harm might defeat a fair-use defence. These plaintiffs, however, had not supplied meaningful evidence of that harm to their own works. Neither ruling gives AI companies unrestricted permission or settles the ethical dispute. They require us to distinguish how books are acquired, how copies are retained and what training does with them. Calling all of this ‘theft’ leaves those questions unanswered.&lt;/p>
&lt;p>Incidentally, the Anthropic ruling provides a rationale for the guillotine. The court upheld the digitisation of purchased copies and rejected the accumulation of pirated books in a permanent library, over which the company later &lt;a href="https://www.anthropiccopyrightsettlement.com/" target="_blank" rel="noopener">agreed to a $1.5 billion settlement&lt;/a>. The digitisation passed in part because each print copy was destroyed in the conversion: the digital copy replaced it and was not shared outside the company. The destruction that so upsets some people helped Anthropic establish that its digitisation programme was fair use. Buying, scanning and discarding offered a lawful route in that case. The guillotine, here, is an instrument of compliance. Buying copies can solve the problem of acquisition. It does not, by itself, give anyone permission to share the resulting corpus.&lt;/p>
&lt;p>There are good reasons to want AI to draw on more languages, more scholarship and more of the world that existed before the internet. Those reasons do not disappear because a company we dislike has found a commercial opportunity in doing so.&lt;/p>
&lt;h2 id="borrowed-indignation">Borrowed indignation&lt;/h2>
&lt;p>Let us return, then, to the Daily Mail. Its article forms part of &lt;a href="https://newsmediauk.org/make-it-fair/" target="_blank" rel="noopener">Make It Fair&lt;/a>, a campaign by newspaper owners and publishers seeking &lt;a href="https://newsmediauk.org/topics/ai-copyright/" target="_blank" rel="noopener">a fully scaled licensing market&lt;/a> for AI. Two convenient substitutions do the work: Big Tech stands for AI, and rights holders stand for British creativity. Readers are invited to defend culture as publishers negotiate the terms of its sale. Scholars sharing the story might ask whose cause their indignation is serving.&lt;/p>
&lt;p>Big Tech and the &lt;a href="https://en.wikipedia.org/wiki/Publishing#Mainstream_publishers" target="_blank" rel="noopener">Big Five&lt;/a> are quarrelling over the rent, but neither objects to the fence. The story of the lone author robbed by a tech giant obscures what expensive licensing deals offer both corporate camps: publishers get paid, and Big Tech gets a market in which its smaller rivals cannot afford to compete. Peter Thiel put Silicon Valley’s preference more bluntly: &lt;a href="https://www.wsj.com/articles/peter-thiel-competition-is-for-losers-1410535536" target="_blank" rel="noopener">‘Competition Is for Losers’&lt;/a>.&lt;/p>
&lt;p>Licensing need not produce that outcome. Affordable, non-exclusive access could help smaller developers. But a system that makes each entrant negotiate costly deals with major rights holders turns purchasing power into a condition of entry. A payment presented as a concession to the cultural industries ends up buying protection against future rivals.&lt;/p>
&lt;h2 id="ai-does-not-belong-to-big-tech">AI does not belong to Big Tech&lt;/h2>
&lt;p>There is a deeper issue. Too much criticism of AI takes its largest commercial suppliers for the whole field. The products dominate the headlines; the research and development beyond those companies disappear from view. Big Tech could hardly ask for a more useful misunderstanding. Its ambition to own the field becomes the premise of the criticism directed against it. The monopoly is not yet secured. Treating it as settled does those companies an enormous favour.&lt;/p>
&lt;p>Open-source AI allows people to build and adapt systems that reflect a greater diversity of languages, cultures and forms of expression. For countries across Europe, Africa and Latin America, as well as South Korea, Japan and India, that is a practical basis for technological and cultural sovereignty: the ability to decide how systems work, which languages they serve and where they are used. Whether to use AI, and for what purposes, should remain a public choice, not a decision dictated by dependence on a foreign supplier.&lt;/p>
&lt;p>But the freedom to modify a system means little without the means to do so. Independent work needs expertise, computing infrastructure and access to training material. That brings us back to the books. If every research team or public institution must negotiate its own publishing deals or assemble its own private library, independence remains a privilege of those who can afford it.&lt;/p>
&lt;h2 id="the-books-are-already-here">The books are already here&lt;/h2>
&lt;p>Libraries and archives have spent generations collecting the material these systems need. Their collections exist because access to knowledge should outlast the commercial interest in selling it. Much has already been scanned: &lt;a href="https://www.hathitrust.org/" target="_blank" rel="noopener">HathiTrust&lt;/a> holds some nineteen million digitised volumes from research libraries, many still in copyright. A library scans a book and keeps it. Nobody there needs a guillotine. HathiTrust is &lt;a href="https://www.hathitrust.org/about/research-center/htrc-transition-faq/" target="_blank" rel="noopener">closing its Research Center at the end of 2026&lt;/a>, but has also announced &lt;a href="https://www.hathitrust.org/press-post/hathitrust-receives-grant-to-deliver-ai-enabled-access-to-digital-books/" target="_blank" rel="noopener">Transparent Books&lt;/a>, a project to provide corpus access for non-commercial AI research and development. Access on those terms would still leave some independent developers outside.&lt;/p>
&lt;p>Owning the books, however, does not give a library the right to distribute their copyrighted contents as a training corpus. &lt;a href="https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng" target="_blank" rel="noopener">EU law&lt;/a> allows research institutions and libraries to mine lawfully accessible works for scientific research, and permits broader mining where rights holders have not reserved their rights. &lt;a href="https://www.legislation.gov.uk/ukpga/1988/48/section/29A" target="_blank" rel="noopener">Britain draws a tighter boundary&lt;/a>: its mining exception covers non-commercial research and restricts the transfer of copies. Neither framework gives libraries a general right to share their copyrighted holdings for others to build on.&lt;/p>
&lt;p>We might also reconsider how long rights holders should control computational uses of books. Pharmaceutical patents offer one point of comparison: their normal twenty-year term is far shorter than copyright’s protection for an author’s lifetime and seventy years beyond. Could books become available for text and data mining without licensing fees after, say, twenty years from first publication, while rights over republication and sale remain protected? Protecting an author’s right to sell a book need not mean controlling every way its contents can be used for generations.&lt;/p>
&lt;p>I argued two years ago in &lt;a href="https://www.thehindu.com/opinion/op-ed/ai-needs-cultural-policies-not-just-regulation/article68469548.ece" target="_blank" rel="noopener">The Hindu&lt;/a> that heritage data should be treated as a public good. Governments should fund libraries and archives to digitise, transcribe and document their collections, and give them the legal authority to make the resulting corpora available for others to build on. Big Tech could help foot the bill through a dedicated levy.&lt;/p>
&lt;p>Public funding should come with published collection policies and access rules, covering academic and commercial work without exclusive deals or privileged access for donors. Access for training need not mean unrestricted redistribution of copyrighted books. Authors should help design the public scheme, including how it addresses consent, remuneration and competition with their work. Free access for users need not preclude publicly funded remuneration for authors. &lt;a href="https://www.bl.uk/services/plr" target="_blank" rel="noopener">Public lending right&lt;/a>, which pays authors from public funds when their books are borrowed, offers one precedent. The fire was never the story. The fence is.&lt;/p></description></item></channel></rss>