Why Do AI Companies Want Old Books?

Older books are becoming valuable AI training material because they provide large quantities of structured, edited and predominantly human-created language. But buying the physical book does not resolve copyright, licensing, provenance, author compensation or cultural preservation. This article explains what is reportedly happening, why it matters, what recent US decisions do and do not establish, why the UK position is different, and what UK businesses should do.
Reports suggest that AI developers and specialist data suppliers are buying large quantities of second-hand books. Some are reportedly scanned by removing the spine so pages can pass rapidly through industrial scanners. The physical book may then be discarded or recycled.
The reason is not nostalgia. Books provide edited language, long-form reasoning, specialist knowledge, consistent structure, material written before generative AI became widespread and data with potentially clearer publication history than anonymous web content. That makes books potentially valuable for training and evaluating language models.
But important questions remain. Who owns the copyright? Was permission obtained? Is the scanning lawful? Is the resulting dataset licensed? Are scarce books being destroyed? Can organisations prove where their AI-training data came from?
The central story
The strange story of AI companies wanting old books is really a story about data quality, copyright and the growing value of verifiably human-created knowledge.
Last checked: 29 July 2026.
What is reportedly happening?
Reports indicate that bulk buyers are purchasing second-hand books, sometimes in large quantities. Orders may range from ordinary used stock to specialist and out-of-print titles. Some buyers may be acting for technology companies or dataset suppliers, but the ultimate customer is not always disclosed. Books may be shipped to scanning facilities where some processes remove the spine, loose pages are fed through high-speed scanners, optical character recognition converts images to text and the resulting datasets may be used to train, fine-tune, test or evaluate AI systems.
| Status | Detail |
|---|---|
| Confirmed | Major AI developers have publicly acknowledged using large quantities of published text in model training, including books |
| Confirmed | Commercial dataset suppliers exist and offer training data derived from books |
| Confirmed | Destructive scanning processes exist and are used in digitisation operations |
| Reported | Booksellers have observed unusual bulk purchasing patterns they attribute partly to AI-related demand |
| Reported | Intermediaries may be purchasing on behalf of technology companies without disclosing the end customer |
| Not publicly known | Complete list of organisations conducting book purchases for AI purposes |
| Not publicly known | Exact quantities involved in any specific operation |
| Not publicly known | Licensing arrangements for specific datasets |
| Not publicly known | Fate of every physical book involved |
A key qualification
The demand is real, but individual supply chains are often difficult to see from the outside.
Not every bulk book purchase is AI-related. The figure should be treated with caution.
AI does not want books. AI developers do.
The language of “AI wanting books” is convenient shorthand for a commercial supply chain. In reality, people and organisations identify useful data, buy or licence books, scan pages, extract text, clean the data, classify it, remove duplication, apply filters, build datasets and train or evaluate models.
The accurate framing
AI does not want your books. AI developers want reliable human-created data.
Anthropomorphic language can obscure commercial decisions, supply chains, copyright choices, procurement, dataset design and human accountability for those decisions.
Why books?
Books can offer professionally edited prose, coherent long-form arguments, specialist vocabulary, narrative structure, historical language, technical knowledge, translated works, multiple writing styles, material not freely available online, text with known author, publisher and publication date, and sustained context over hundreds of pages.
Compare this with internet data, which may contain duplication, spam, scraped fragments, weak attribution, misinformation, SEO-generated text, low-quality automated content, machine translation, AI-generated material, malicious data and missing context.
A useful distinction
A published book is not automatically true — but it is often more structured and attributable than an anonymous web page.
Books can also contain outdated information, historical prejudice, factual errors, regional assumptions, incomplete perspectives, copyrighted material, personal information and transcription errors. Older text may reflect social attitudes that have since changed. These limitations do not eliminate the training value, but they do require careful filtering and assessment.
Why older books?
Books produced before widespread generative-AI adoption are more likely to contain human-created text that has not itself been generated from previous AI systems. However, using any specific date as an absolute dividing line is too simple. Generative AI existed before 2022. Human writing continues after 2022. Publication dates are only a rough signal. Some newer books may contain AI-generated content. Some older digitised texts may contain OCR errors that introduce machine-readable problems. Provenance matters more than date alone.
The classification challenge
“Pre-AI” is a useful shorthand, not a perfect classification.
Developers may value older material because it can help reduce the risk of training new models primarily on outputs created by earlier models. As more internet content is generated by AI tools, finding large quantities of text that demonstrably predates the generative-AI era becomes more commercially valuable.
The human-data premium
Clearly attributable human-created material may become more commercially valuable as synthetic content becomes more common. Potentially valuable characteristics include: named author, known publication date, editorial process, identifiable publisher, documented edition, reliable subject classification, verified language, known rights holder, established source and limited AI-generated contamination.
A wider implication
In an internet full of generated material, proof that information came from a human may itself become valuable metadata.
This has implications for publishers, archives, professional associations, universities, technical authors, specialist businesses, research organisations, long-established websites and owners of proprietary document collections. The value may lie not only in the content itself but in the demonstrable provenance that accompanies it.
The bigger picture
AI may appear digital, but its capabilities are still built from human-created knowledge.
How destructive scanning works
- 1The book is acquired
- 2Its binding or spine may be removed to allow individual pages to be separated
- 3Loose pages are fed through a high-speed scanner
- 4Images of the pages are captured
- 5Optical character recognition converts images into machine-readable text
- 6Text is checked, cleaned and structured to remove common OCR errors
- 7Metadata may be added: author, title, publisher, date, language and subject
- 8The resulting files enter a dataset or document archive
Destructive scanning is used because it is faster than page-by-page photography, has lower labour cost, is easier to automate, produces cleaner page alignment and supports higher scanning volume. The disadvantages are significant: the physical volume may be permanently damaged or destroyed, annotations or inserts may be lost, binding and production history may disappear, provenance may be weakened, unique copies may be destroyed and OCR may introduce errors.
Does the book have to be destroyed?
No. Non-destructive alternatives include overhead book scanners, cradle scanners, specialist archive photography, manual page turning and library digitisation systems. These methods may be slower, more expensive, labour-intensive and difficult at very large scale.
An important clarification
Destroying the book is a commercial or procedural choice, not an inherent requirement of artificial intelligence.
Destruction may arise from the chosen scanning method, cost reduction, throughput targets, storage decisions or the low resale value of ordinary stock. The decision to destroy, retain or donate the physical object is separate from the technical requirements of the digitisation process itself.
Buying a book is not buying its copyright
The fundamental principle
Buying the paper does not automatically buy the right to reproduce the words.
When someone buys a printed book, they normally acquire that particular physical copy. They do not automatically acquire the author’s copyright, the publisher’s rights, permission to reproduce the entire work, permission to distribute a digital copy, permission to create a commercial dataset, permission to use the work in every jurisdiction or permission to generate substantially similar material.
| What was acquired? | What it normally allows | What it does not automatically allow |
|---|---|---|
| Physical copy | Reading, keeping, reselling, gifting, disposing of that copy | Mass reproduction, digital distribution, commercial dataset creation, AI training |
| Copyright or licence | Specified copying, digitisation, dataset use, model training, defined commercial uses (all subject to terms, territory, duration and permitted purpose) | Uses beyond what is defined in the licence or permitted by applicable law |
The legal reality
Ownership of an object and ownership of the intellectual property within it are not the same thing.
Physical ownership, copyright ownership, permission to digitise, permission to retain a digital copy, permission to use that copy for AI development and permission to produce commercial outputs are separate questions. Each requires separate assessment.
Legal outcomes depend on jurisdiction, the specific use, how the material was obtained, applicable licences, statutory exceptions and the facts of the individual case. This article provides general business information and is not legal advice.
What have US courts actually decided?
Several significant cases in US federal courts have examined the relationship between AI model training and copyright. These cases have involved questions including whether using copyrighted works for model training constitutes fair use, how copies were acquired, whether source copies were from unlicensed or pirated sources, whether lawfully purchased physical books differ from unlicensed digital copies, whether intermediate copies were retained, whether the use was transformative, market effects and resulting outputs.
US fair-use analysis is fact-specific: courts weigh the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used and the effect on the market for the original. Results in individual cases do not create universal rules. Decisions may be appealed, settled, or limited to specific facts. The procedural stage of a case — whether it is a motion to dismiss, summary judgment, full trial or appeal — affects how much weight it should be given. Some cases result in settlements rather than final judgments that create binding precedent.
The limits of a US ruling
A court ruling about one developer, one dataset and one set of facts is not a global licence for AI training.
Why the UK position is different
UK copyright law uses fair dealing rather than the broad US fair-use doctrine. Fair dealing is a more specific set of exceptions with defined conditions, including the purpose of the use, whether the source is acknowledged and whether the amount used is justified. Text-and-data-mining exceptions in UK law have specific conditions. Commercial and non-commercial uses may be treated differently. Contractual and licensing arrangements may restrict or extend what statutory law would otherwise permit.
The UK Government has examined the intersection of AI and copyright, including through its response to consultation and in policy developments from the Intellectual Property Office. That landscape continues to develop. Uncertainty remains around transparency obligations, licensing, enforcement mechanisms and the position of cross-border model development.
A clear caution for UK organisations
For a UK organisation, “a US court allowed it” is not an adequate copyright assessment.
- No US judgment automatically binds a UK court
- UK organisations should not rely on US headlines for copyright decisions
- Cross-border model development creates complex jurisdiction questions
- AI services may be developed, trained, hosted and used under different legal frameworks simultaneously
- Businesses should obtain specialist legal advice where material risk exists
What is training-data provenance?
Provenance is the documented history of where information came from, how it was obtained and what may lawfully be done with it. For each dataset, useful records include: source, creator, rights holder, publication date, acquisition method, licence, permitted use, territory, restrictions, retention policy, quality checks, opt-outs, deletion requests and responsible owner.
Provenance matters because it supports legality assessment, model-quality evaluation, bias assessment, accuracy verification, supplier assurance, dispute handling, regulatory compliance, customer trust and the ability to remove or retrain if a problem is identified.
The provenance principle
If a supplier cannot explain where its training data came from, it cannot fully explain the risk contained within its model.
The quality principle
The quality of an AI system begins long before anyone types a prompt.
AI-generated data and model degradation
Researchers have raised concerns about repeatedly training models on synthetic outputs. Possible risks include amplification of errors, reduced diversity, reinforcement of common patterns, loss of rare information, feedback loops, homogenised language, declining source attribution and greater difficulty distinguishing human and machine material.
Model collapse is a term used for degradation that may occur when successive systems are trained heavily on generated material and lose parts of the original data distribution. Research in this area is ongoing and the extent of the risk under different training approaches is actively debated.
The precise concern
The problem is not simply synthetic data. It is synthetic data whose origin, quality and influence are unknown.
Well-designed synthetic data may be useful when it is deliberately generated, accurately labelled, quality-controlled, balanced with suitable real-world data and used for specific tasks. The concern is uncontrolled synthetic contamination of datasets where its presence and extent are not tracked.
Are rare books at risk?
Concerns may arise when bulk purchasing includes scarce editions, out-of-print technical works, local histories, minority-language publications, annotated copies with valuable marginalia, limited academic publications, small-print-run books, privately published works or books not preserved elsewhere. An ordinary second-hand paperback held in thousands of library copies is a different matter from a unique annotated copy or the last accessible copy of a local historical work.
Questions worth asking in any large-scale digitisation project include: Was scarcity checked? Is another preserved copy available? Were annotations recorded? Was the edition documented before scanning? Could a non-destructive method have been used? Should libraries or archives have been consulted? Is the digital file preserved responsibly? Who controls future access?
The preservation concern
Digitising the text may preserve the words while still destroying information carried by the physical object.
Authors, publishers and licensing
Possible licensing approaches include direct licences between developer and rights holder, publisher licences covering catalogues, collective licences through collecting societies, opt-in schemes with revenue sharing, per-title payments, dataset subscriptions, research-only licences, limited-use licences and model-specific licences with output restrictions or attribution requirements.
Potential benefits of licensing include permission clarity, compensation for creators, transparency, higher-quality datasets with known provenance, dispute reduction and author choice. Challenges include identifying rights holders, split rights between author and publisher, older contracts that did not anticipate AI use, orphan works, territorial rights, administrative cost, pricing, monitoring compliance, collective representation structures, withdrawal after training has occurred and measuring individual contribution.
What this means for ordinary businesses
Most SMEs are not training foundation models. However, businesses may buy AI software, commission custom models, fine-tune systems, create retrieval databases, upload document libraries, licence generated content, publish AI-assisted material, use AI outputs commercially or rely on suppliers’ legal assurances.
Potential risks include supplier copyright claims, outputs resembling protected material, unclear indemnities, inability to identify sources, confidential material entering training, contractual restrictions, reputational damage, loss of proprietary content, regulatory scrutiny and customer concerns.
The business lesson
You may not control the original training dataset, but you can control which supplier you trust and what evidence you require.
Questions to ask about AI training data
- 1What types of data were used to train the system?
- 2Were books, articles, images or other copyrighted works included?
- 3How was the material obtained?
- 4Was it purchased, licensed, scraped or supplied by a third party?
- 5Can the supplier document provenance?
- 6Which jurisdictions governed acquisition and training?
- 7Which copyright exceptions or licences are relied upon?
- 8Were pirated or unauthorised sources excluded?
- 9How are rights-holder objections handled?
- 10Can material be removed from future training?
- 11Is customer data used for model training?
- 12Is customer data used to improve the service?
- 13Can training use be disabled?
- 14Are prompts retained?
- 15Are uploaded documents retained?
- 16Are outputs checked for memorised content?
- 17What measures reduce verbatim reproduction?
- 18Has the dataset been checked for personal data?
- 19Has the dataset been checked for bias?
- 20Has synthetic material been identified?
- 21How is synthetic data labelled?
- 22What model cards or data documentation are available?
- 23Have independent copyright or data-governance reviews been completed?
- 24What contractual indemnity is offered?
- 25Who bears the cost of an intellectual-property claim?
- 26What happens if a dataset is later found to contain unlawful material?
- 27How are model versions and dataset changes documented?
- 28Will customers be notified of material changes?
- 29Can the supplier identify authoritative sources used for specific business functions?
- 30What audit evidence is available?
Protecting your own business content
Organisations that own reports, articles, manuals, training materials, photographs, product descriptions, research, databases, designs, customer documentation, knowledge bases or software documentation should consider the following:
- Identify valuable intellectual property and record ownership clearly
- Review employment and contractor agreements to confirm who owns created content
- Add appropriate copyright notices to published material
- Define permitted AI use in internal policies
- Review website terms and API terms to address crawling and training use
- Control bulk downloads and monitor unusual scraping activity
- Maintain original files with creation dates as evidence of authorship
- Use access controls where appropriate
- Consider licensing opportunities for high-quality proprietary content
- Document authorised datasets for any custom AI projects
- Review AI crawler controls and consider robots.txt configuration
- Seek legal advice for valuable or disputed material
A realistic caveat
Technical controls, contractual terms and copyright records serve different purposes and work best together.
Robots.txt instructions may not be followed by all crawlers. Metadata does not guarantee enforcement. Copyright notices do not prevent infringement. Crawler blocking does not remove copies already made. An opt-out mechanism works only if the relevant supplier honours it. Legal advice remains necessary where material risk exists.
Immediate actions for businesses
- 1Identify your AI suppliers — record every generative-AI service in use
- 2Review contracts — check training, retention, intellectual-property and indemnity terms
- 3Ask about data provenance — request clear supplier documentation
- 4Disable customer-data training where appropriate — use available enterprise controls
- 5Control document uploads — prevent staff uploading material without authority
- 6Check outputs before publication — look for copied passages, false attribution and unsupported claims
- 7Protect valuable content — record ownership and review access
- 8Assess custom AI projects — document the source and licence of every dataset
- 9Review third-party datasets — do not assume “publicly available” means freely reusable
- 10Assign ownership — give AI procurement a named business and technical owner
- 11Monitor legal developments — copyright and AI policy continues to change
- 12Record decisions — keep evidence supporting the chosen supplier and use case
The procurement principle
Do not buy an AI system solely on what it can produce; assess what it was built from and what contractual protection sits behind it.
Warning signs
Investigate further if an AI supplier:
- Refuses to discuss training-data sources
- Describes all internet content as free to use
- Relies on a single foreign court headline
- Confuses fair use with fair dealing
- Cannot identify dataset suppliers
- Offers no model card or dataset documentation
- Gives vague answers about customer-data training
- Claims copyright risk is impossible
- Provides no intellectual-property terms
- Excludes all liability
- Offers no process for rights-holder complaints
- Cannot explain data retention
- Cannot explain model updates
- Has no approach to memorised output
- Uses undocumented third-party datasets
- Cannot distinguish licensed, public-domain and scraped material
- Cannot confirm how personal data is handled
- Makes sweeping claims that its model owns all generated outputs
- Will not explain whether customer documents improve the general model
Not every supplier must disclose commercially sensitive dataset details publicly. The question is whether sufficient assurance is available to the customer.
Practical business implications
Training data is part of supplier risk
The model’s origin matters as much as its features. Evaluating an AI system only on what it produces, without examining what it was built from, leaves significant risk unassessed.
Publicly accessible does not always mean free to reuse
Access and permission are different concepts. A book available in a second-hand shop is accessible but the copyright belongs to the author and publisher. A website indexed by search engines is accessible but subject to its terms of service and applicable copyright law.
Human-created information may become more valuable
Reliable provenance can become a competitive asset. Organisations with large collections of high-quality, clearly attributable, human-created content may find that content has training-data value beyond its original purpose.
Copyright risk crosses borders
AI services may be developed, trained, hosted and used in different jurisdictions simultaneously. A model trained in one country, offered by a company based in another and used by a UK business creates a complex multi-jurisdictional copyright question. UK businesses should not assume foreign legal positions resolve their domestic obligations.
Output ownership is only part of the question
Businesses must also consider the material used to build the system. Contracts often address ownership of outputs without clearly addressing the copyright position of the training data underlying those outputs.
AI procurement needs legal and information-governance input
This should not be left solely to technical teams. Copyright, data protection, contractual indemnity and regulatory compliance all require appropriate expertise.
Small businesses also own valuable data
Manuals, customer knowledge, specialist content and proprietary research may have commercial training value. Recording ownership and considering licensing terms is relevant even for organisations that have not previously thought of themselves as content businesses.
The governance principle
AI governance begins with inputs, not merely outputs.
The IT Club View
The image of AI companies scouring second-hand shops for old books feels strange because artificial intelligence is normally presented as entirely digital. In reality, the latest technology still depends heavily on human-created material accumulated over decades or centuries.
Older books may be attractive because they offer structured writing, specialist knowledge, known publication histories and material less likely to have been generated by AI. That does not make every use lawful, ethical or culturally harmless. The central questions remain: Was the material obtained legitimately? Was the use authorised? Can the supplier prove provenance? Were scarce physical works protected? Were creators recognised or compensated? Can customers understand the resulting risk?
A broader point
The future of AI may depend partly on preserving — and fairly valuing — the human knowledge that came before it.
IT Club is not arguing that AI development should stop, that books should never be digitised, that synthetic data has no value, or that all scanning is destructive. Instead, the case is for transparency, provenance, lawful acquisition, proportionate licensing, preservation, supplier assurance and accountable procurement.
The IT Club View
Better AI should not require businesses to stop asking where its knowledge came from.
Reviewing your AI suppliers?
Do you know what your AI supplier’s model was built from? Use our AI Training Data Supplier Checklist to review provenance, copyright, licensing, customer-data use, retention, output protection and contractual responsibility.
Related business questions
Why do AI companies want old books?
Older books provide large quantities of structured, professionally edited and predominantly human-created text. As more internet content is generated by AI tools, text that demonstrably predates the generative-AI era becomes more commercially valuable for training and evaluating language models. Books also offer sustained context, specialist vocabulary and known publication histories that are harder to establish for anonymous web content.
Are AI companies buying second-hand books?
Reports from booksellers and journalistic investigations suggest that bulk purchasing for AI-related purposes is occurring, though the full extent is not publicly known. Some purchases may be made through intermediaries without disclosing the end customer. Not every bulk book purchase is AI-related, and the scale of any specific operation is often difficult to verify from outside the supply chain.
Why are books useful for AI training?
Books can provide professionally edited prose, long-form arguments, specialist knowledge, consistent structure, material not freely available online and text with a known author and publication date. This contrasts with internet data, which often contains duplication, spam, AI-generated material and weak attribution. Books are not automatically accurate or unbiased, but they are often more structured and attributable than anonymous web content.
Why is pre-AI material valuable?
Material produced before widespread generative AI is more likely to be predominantly human-created rather than itself generated from earlier AI systems. This helps reduce the risk of training new models primarily on synthetic outputs. As more online content is machine-generated, text with a reliable human provenance becomes more commercially significant. Publication date is a rough signal rather than a perfect classification.
What is human-created training data?
Human-created training data is material written, edited and published by people rather than generated by software. It includes books, articles, academic papers, technical documentation, transcripts and original web writing. As synthetic content becomes more prevalent, material with demonstrable human authorship and editorial process may be considered higher quality for certain training purposes.
What is synthetic-data contamination?
Synthetic-data contamination occurs when AI-generated content enters a dataset without being identified as such. If a model is then trained on that data, it may reinforce patterns from earlier generated content rather than learning from original human expression. This can reduce diversity, amplify errors and make it harder to distinguish human and machine material in future datasets.
What is model collapse?
Model collapse is a term used for degradation that may occur when successive AI systems are trained heavily on generated material. The concern is that the model may lose parts of the original data distribution, producing more homogenised output with reduced diversity. Research is ongoing and the conditions under which this occurs at significant scale continue to be studied. Deliberately generated, well-labelled and quality-controlled synthetic data is not the same risk.
What is destructive book scanning?
Destructive book scanning involves removing the binding or spine of a physical book so individual pages can be fed through a high-speed scanner. This allows faster, cheaper and more automated digitisation than non-destructive alternatives, but the physical volume is permanently damaged or destroyed in the process. Annotations, marginalia and binding history are lost. Not all scanning operations use this method.
Are scanned books always destroyed?
No. Non-destructive alternatives include overhead book scanners, cradle scanners and specialist archive photography. These preserve the physical object but are slower, more expensive and harder to scale. Destruction is a commercial or procedural choice, not an inherent requirement of digitisation. Whether a specific operation destroys its source materials depends on the method chosen and the commercial decisions of the organisation involved.
Does buying a book transfer copyright?
No. Purchasing a physical book normally transfers ownership of that copy only. It does not transfer the author’s copyright, the publisher’s rights or any permission to reproduce, distribute or commercially exploit the text. The copyright remains with the rights holder regardless of who owns the physical object.
Can a company scan a book it owns?
Scanning a book you physically own may or may not be lawful, depending on the jurisdiction, the purpose, the use made of the resulting copy, applicable copyright exceptions and any licence held. In the UK, text-and-data-mining exceptions have specific conditions. Personal use, research and commercial use are treated differently. Simply owning the physical copy does not automatically permit scanning, retaining, distributing or commercially using a digital copy.
Is using books for AI training legal?
It depends on the jurisdiction, how the material was obtained, whether a licence or applicable exception exists, the purpose of the use and the specific facts. Legality cannot be assumed from the fact that a book was lawfully purchased. UK and US law treat relevant exceptions differently. Active litigation continues in multiple jurisdictions. This article provides general business information and is not legal advice. Businesses with material exposure should obtain specialist legal advice.
Does US fair use apply in the UK?
No. US fair use and UK fair dealing are different legal doctrines. Fair use in the US involves a flexible, case-by-case balancing test with four statutory factors. UK fair dealing is a more specific set of exceptions with defined conditions. A conclusion reached by a US court under fair use does not automatically apply in the UK. UK organisations should not rely on US legal outcomes when assessing their position under UK copyright law.
What is the UK position on AI and copyright?
The UK has examined AI and copyright through its Intellectual Property Office and Government consultation. Text-and-data-mining exceptions in the Copyright, Designs and Patents Act 1988 have specific conditions regarding commercial use. The UK Government’s position has developed over time and continues to be reviewed. Businesses should check current official guidance rather than relying on older commentary, and should obtain legal advice where material risk exists.
What is training-data provenance?
Training-data provenance is the documented history of where AI training material came from, how it was obtained, what rights apply and what may lawfully be done with it. Good provenance records include source, creator, rights holder, acquisition method, licence, permitted use, jurisdiction, retention policy and quality checks. Provenance matters for legality, model quality, bias assessment and the ability to respond to rights-holder requests.
Can AI reproduce material from its training data?
Research has shown that large language models can sometimes reproduce passages from their training data verbatim, particularly for material that appeared frequently. The conditions under which this occurs, how it is detected and what controls are effective remain active areas of research. Responsible AI systems typically implement measures to reduce verbatim reproduction, but complete prevention is technically challenging.
Are rare books being destroyed?
There are legitimate concerns that bulk scanning operations may not adequately assess whether individual volumes are scarce or culturally significant before using destructive methods. An ordinary second-hand paperback available in many libraries is different from a rare annotated edition or a locally significant work. Whether specific rare books have been destroyed is not publicly known in most cases. The concerns are proportionate rather than universal.
Can authors licence books for AI training?
Yes, where they hold the relevant rights. Rights in a published book may be shared between author and publisher, and older contracts may not address AI training. Collective licensing through authors’ organisations or collecting societies is one possible mechanism. Various licensing models are being explored, including opt-in schemes, revenue sharing and per-title payments. The practical and administrative challenges are significant, particularly for older catalogues.
What should businesses ask an AI supplier?
Key questions include: What data was used to train the system? How was it obtained? What licences or legal exceptions apply? Which jurisdictions governed acquisition and training? Is customer data used for training? Can training use be disabled? What contractual indemnity is provided? How are rights-holder objections handled? Our AI Training Data Supplier Checklist provides a complete review structure.
Is customer data used to train AI?
It depends on the supplier and the terms of the contract. Some AI services use customer interactions to improve future model versions by default; others offer enterprise controls to opt out. Some distinguish between base-model training and service improvement. Prompts, uploaded documents and outputs may be retained and used differently. Review the supplier’s privacy policy, data-processing agreement and contract terms carefully, and use available enterprise controls.
How can a business protect its own content?
Practical steps include recording ownership of valuable content, reviewing employment and contractor agreements, adding copyright notices, reviewing website and API terms, controlling bulk downloads, monitoring unusual scraping, retaining original files with creation dates, considering licensing opportunities and seeking legal advice for material exposure. Technical controls, contractual terms and copyright records serve different purposes and work best together.
Does robots.txt stop AI training?
Not reliably. Robots.txt is a technical convention that some crawlers respect and others do not. It is not legally binding and provides no guarantee that content will not be collected or used. It is one layer of control, but it should not be treated as a sufficient legal protection on its own. Contractual terms and copyright records remain necessary.
Who is responsible if AI output infringes copyright?
This depends on the applicable law, the terms of the contract with the AI supplier and the specific facts. In many supplier agreements, the customer bears some responsibility for how outputs are used. Some suppliers offer intellectual-property indemnities, either as standard or as part of enterprise contracts. The extent of those indemnities varies significantly. Businesses should review indemnity terms before commercial use of AI-generated content.
Administrator Technical Note
This section is intended for IT administrators, data-governance leads, legal and compliance teams, and technical staff responsible for AI procurement and dataset management.
Training-data lifecycle
A responsible training-data lifecycle covers: source identification, rights and licence assessment, acquisition, digitisation or ingestion, OCR, cleaning, normalisation, language detection, classification, deduplication, personal-data filtering, quality filtering, safety filtering, provenance recording, dataset versioning, training or fine-tuning, evaluation, retention or deletion, rights-holder request handling and audit.
Every stage may introduce errors, missing metadata, legal risk, bias, security risk or quality degradation. Treating dataset governance as a one-time activity at acquisition rather than an ongoing process creates compounding risk.
OCR and data-quality risks
OCR processes may introduce missing characters, merged words, incorrect punctuation, page-order errors, footnotes inserted into body text, repeated headers, incorrectly read columns, corrupted tables, lost mathematical notation, misspelled names, language misidentification and detached captions. A clean-looking scan may still contain machine-readable errors at enormous scale. Quality filtering and sampling are necessary rather than optional.
Deduplication
Duplicates in book-derived datasets may arise from multiple editions, reprints, ebook versions, web copies, quotations, anthologies, translations, mirrors and OCR variants of the same text. Poor deduplication can overweight particular works, increase memorisation, distort model behaviour, amplify bias and make evaluation unreliable.
Provenance architecture
Recommended provenance controls include immutable source identifiers, licence records, content hashes, acquisition timestamps, supplier records, version-controlled manifests, transformation logs, dataset lineage tracking, model-to-dataset mappings, deletion markers, rights-holder request records, audit trails, access controls and named accountability. Provenance records should remain available even if source content is later removed.
Dataset inventory fields
Recommended fields: dataset name, version, owner, supplier, source categories, acquisition date, acquisition method, copyright status, licence, permitted purpose, prohibited purpose, jurisdiction, rights holder, publication date, creator, personal-data status, sensitive-data status, synthetic-data status, OCR method, quality checks, deduplication method, bias review, retention, deletion process, opt-out process, model versions trained on this data, downstream recipients, last review, incidents and decision.
Customer data and retrieval systems
Foundation-model training, fine-tuning, retrieval-augmented generation, prompt context input and logging or service improvement involve different legal implications, retention risks, confidentiality risks, deletion options and contractual controls.
Uploading a document to an AI service does not always mean it trains the base model — but the contract must confirm what actually happens.
Memorisation and output testing
Controls relevant to memorisation and verbatim reproduction include: verbatim-reproduction testing, long-matching-sequence detection, source attribution where appropriate, canary testing, prompt-based extraction testing, testing against known copyrighted benchmark material, repeated-phrase detection, model-version comparison, escalation procedures and suppression mechanisms.
Data security
Book-derived datasets may still contain personal data, addresses, case studies, correspondence, medical details, historical records, confidential material, security-sensitive information, outdated credentials or harmful instructions. Published material is not automatically risk-free. Data classification, personal-data review, access control, encryption, retention limits, monitoring, incident response and removal procedures are all relevant.
Supplier contract controls
Recommended contract review areas: training-data warranty, lawful acquisition warranty, non-infringement language, customer-data training terms, confidentiality, data retention, output ownership, third-party claims, indemnity, liability caps, notification obligations, cooperation requirements, model-change provisions, subcontractor terms, dataset supplier identification, audit evidence, termination, deletion, governing law and jurisdiction. Legal review is advisable for high-value or high-risk AI contracts.
Data licensing controls
Machine-readable licence fields can support automated compliance checking. Key licence fields include: permitted training use, evaluation-only use, commercial restrictions, territory, term, sublicensing rights, derivative dataset permissions, model output conditions, retention after termination, deletion obligations, audit rights, attribution requirements, compensation arrangements, opt-out process and dispute procedure. Metadata alone does not determine legal rights.
Operational Heartbeat
AI training-data risk changes as models are updated, suppliers change datasets, court decisions develop, government policy changes, licences expire, rights holders object, datasets are acquired, customer-data terms change, new model versions are released, synthetic data increases in the training pool, suppliers merge, subcontractors change, intellectual-property claims arise, retention periods expire and business use expands.
A recurring review should check: AI supplier inventory, current model versions, training-data statements, customer-data settings, contract changes, dataset provenance, licences, supplier litigation, regulatory developments, intellectual-property complaints, output testing, uploaded document controls, staff use, model updates, named owners, corrective actions and next review date.
Operational Heartbeat
AI training data needs an operational heartbeat: sources, licences, supplier terms, model versions and legal developments should be reviewed rather than assumed to remain unchanged.
Plain-English Takeaway
AI companies are interested in older books because they contain large quantities of structured, predominantly human-created information. Some books may be scanned destructively, but buying a physical copy does not automatically transfer the copyright or settle whether the text can be used for AI training. Businesses should ask AI suppliers where their training data came from, what licences or legal grounds apply and what protection is provided if that material is later challenged.
Need the practical steps?
A short, instruction-led version of this topic is available in the Knowledge Centre.
View the Knowledge Centre GuideRelated Articles
Would You Let an AI Coach Your Employees?
AI coaching platforms can now let employees practise workplace conversations with interactive avatars and receive automated scoring and feedback. That could make training more accessible and repeatable. The question is whether the practice room quietly becomes a performance monitoring system.
Read articleDid an AI Really Escape and Launch a Cyberattack?
A reported AI security incident has attracted dramatic headlines. But the most important lesson is not about a machine gaining consciousness. It is about what happens when a capable system is given powerful tools, broad access and insufficient containment.
Read articleHow to Find Every Photo on Your Windows PC
Photographs often become scattered across a Windows PC — saved in Pictures, Downloads, Desktop, OneDrive, project folders and places you have long forgotten. File Explorer can search across the entire computer and display image files from many locations in one results view. Here is how to do it safely.
Read article