Stripped to the Spine: The King's Library and the Limits of AI's Archive

Stripped to the Spine: The King's Library and the Limits of AI's Archive

I.

The King's Library at the British Library contains thousands of books from the collection of one of the most prominent bibliophiles in English history: King George III. While some AI developers might be too stuck on the idea of stripping books down to their spines for data, I'd rather read some of them. The King's Library books would be an example of the perspectives that were prevalent at the time (whatever was thought to be "true" from a mainly British/Eurocentric point of view) without me knowing which titles are housed in this library.

Not everything that is true is written. Not everything is written.

II.

The King's Library was George III's. He began curating it around 1760, and by his death in 1820 it held over 65,000 printed volumes. His son, George IV, gave it to the nation in 1823 rather than keep it at Buckingham Palace, on the condition that it be kept "entire, and separate".

It sat in the British Museum for over 170 years before moving into the British Library. It's now housed in a six-storey glass tower, the first thing most visitors see. The tower was designed by an architect. The clearness of the glass lets a viewer see that these books exist at all. The bindings sit in a climate-controlled environment. It preserves and presents at the same time.

I'm registered at the British Library and have a Reader Pass. I still wasn't sure, standing near the tower, whether anyone gets to hold these books, or whether they've become something just to look at. I've since discovered that with a Reader Pass, you can order items by shelf mark to view in a Reading Room. Some books in the Kings Library go back to the 1400s. These rare books need to be requested in advance. The glass tower isn't keeping people out. You just have to know: that the system exists; how to search a shelf mark; this is a thing you're allowed to ask for. Privilege through procedure.

How to access the King's Library

III.

I keep wanting to call the collection Eurocentric, and I have to be honest about what I actually know. I haven't read these books and I don't know their exact titles. I can't claim their contents are Eurocentric as that would be inventing evidence I don't have. What I can claim (because it's documented) is the conditions of the libraries assembly. George III set budgets for the collection, his purchasing agents visited various European cities. A founding part of the Kings Library was purchasing the library of Joseph Smith (the British Consul at Venice). None of that tells me what any single volume argues. All of it tells me what kind of world could produce a "universal library" in the first place: European print markets, colonial trade routes, a monarch's disposable income, and the scholarly networks that decided what counted as worth buying. The books I'd most want to find are the texts his own political and legal education ran on. Politics, philosophy, and economics were relational for him long before universities formalised that relationship into a degree.

What complicates the story I want to tell is that George didn't just absorb this material uncritically. Historians researching the Georgian Papers have found that he drew critical reflections on slavery and the slave trade directly from reading Montesquieu, in the late 1750s, before he had any power to act on what he'd written. The canon I want to call Eurocentric contained its own internal dissent. A future king read it and pushed back against parts of it, privately. The right critical material existed and still didn't survive into the story anyone tells about this library now. Beyond the framing as simply "keen bibliophile," the young monarch quietly argued with his own philosophers about the practice that funded his empire, conflicted between "private morality and public morality".

V.

I think about what these books are, physically, right now. Leather covers as a status symbol independent of what was inside. They're old enough now that the lignin in the paper has been breaking down for two centuries, releasing Vanillin as it does (the same compound found in vanilla beans). The King's Library almost certainly smells sweet. An archive of imperial and philosophical self-fashioning, quietly decaying. The material doesn't care what it argued, what's on the pages, it just ages.

VI.

The books an AI developer might strip for data have survived through a largely undisturbed lineage in European and American print culture. But that continuity doesn't always follow a repeatable pattern, it's more fragile than it looks. In the 2023–24 US school year, PEN America recorded 10,046 instances of individual books being banned from schools (more than triple the year before). The following year saw 6,870 bans, still well above the pre-2023 baseline. A print archive that looks stable from the outside can be politically disrupted within a single academic year.

Compare that to Babylonian clay tablets, thousands of years older, that have outlasted the empires that made them, the languages inscribed on them falling out of use, and the rooms that once held them no-longer exist. The British Museum have held more clay tablets (about 130,000) than the Iraqi Museum does (of the cuneiform tablets that are 'known' to exist). Cuneiform survived not because anyone protected it as heritage, but because fired clay resists fire, damp, and most forms of censorship in a way paper never has. What was recorded on a tablet was recorded for that time, for that place - it may or may not hold anything relevant now. It only becomes relevant if it forces a question I hadn't thought to ask. That's the same thinking I've been applying to the King's Library throughout this piece: not what's on the shelves, but what conditions produced them, and what that forces me to ask that I wouldn't have asked otherwise. The King's Library, and whatever an AI developer scrapes from libraries like it, is drawing from a medium and a moment more fragile and more politically dependent than most humans would like to admit.

What survives to be catalogued, digitised, and eventually scraped is shaped by nation-states deciding what counts as worth preserving in the first place. The acquisition logic already at work in George III's collection now operates through national archives rather than a monarch's budget. Digitisation adds a second layer of selection on top of the first: research on AI and cultural heritage bias shows it's not just what gets digitised, but how - the tools and priorities used to do the digitising shape the record all over again, before an AI developer ever touches it.

What's available to strip is shaped by copyright law rather than editorial judgement: in the UK, literary works published before 1989 remain protected until 31 December 2039 regardless of how long ago the author died. After that date, government policy is to drop the rule and revert to the standard term of 70 years following the authors passing. The boundary of what's freely available to an AI developer is a legal accident of timing, not a judgement about what deserves to be read.

VII.

What's determined "true" in the world was never universal: not in theology, politics or economics. Some things get recorded because the instruments to measure them exist, and those instruments are unevenly distributed: there are thousands of synchrotron beamlines worldwide, the instruments materials scientists rely on to determine what something is actually made of at a structural level, concentrated almost entirely in the Global North. Africa remains the only inhabited continent without one, and Latin America has just a single facility. Whoever has access to that machinery gets to generate "facts" about a material. Whoever doesn't, doesn't - not because the material is any less real, but because nobody nearby has access to the equipment to measure it.

And some things were never going to be measured at all. Family stories mostly stay oral, especially for people considered at the lower rungs of a society, or wherever publishing itself was suppressed or simply never made available. Oral history exists as its own discipline because written archives failed to hold the lives of "the under-classes, the unprivileged, and the defeated"(p5).


Language drifts the same way. Dominant languages change shape generation by generation, absorbing and shedding vocabulary as their speakers move through history. Smaller languages survive where they're protected by isolation rather than institutional care - in mountain valleys (like Romansh), in small communities, wherever a dominant tongue hasn't yet reached.

Whatever an AI developer trains on will overwhelmingly reflect whichever languages were already dominant enough, and printed enough, to be scraped.

Whoever does the training matters. Women make up a fifth of the UK AI workforce, a proportion that has fallen since 2020. A narrow demographic curating from an already narrowly preserved selection of books is not a new problem.

After reading Edward Said's 'Orientalism' (many years ago) I bought a copy of Lord Cromer's 'Modern Egypt' for myself. Britain's Consul-General in Egypt for over two decades wrote with total confidence about "the Oriental mind" from a position of near-total authority over how Egypt was described, governed, and understood. His version of what Egypt and Egyptians were, knowledge and power fused into one voice. Lord Cromer first published this work in 1908, decades after the Kings Library was formed, into a print culture that could produce far more copies than George III's agents ever could. That's likely why a second-hand copy was still findable for me to buy at all.


A text-based quote from Lord Cromer's 'Modern Egypt', printed in black on a white background. Quoting an observation about the 'Oriental mind' by Sir Alfred Lyall. "Sir Alfred Lyall once said to me : " Accuracy is abhorrent to the Oriental mind. Every Anglo-Indian official should always remember that maxim." Want of accuracy, which easily degenerates into untruthfulness, is, in fact, the main characteristic of the Oriental mind."
Quote from page 146 of Lord Cromer's 'Modern Egypt'


'Modern Egypt' was published over a hundred years ago. Seventy-five years ago, Britain's empire was in retreat, but still standing. If AI developers strip books down to their spines and only find a familiar pattern....

Data gaps aren't neutral, they're designed. So is what the King's Library holds: shaped by budgets, by agents told who not to outbid, by what a print market in the 1700s could even offer for sale, by whose voice was worth writing down and whose wasn't. The King's Library is a well-documented example of a much larger and much less documented pattern: which archives get funded, measured, protected and which don't. I still don't know exactly what's on those shelves. But I understand more clearly now how that gap was designed. If AI developers are going to strip books down to their spines, they might want to consider expanding their worldview first. It might surprise them.

VIII.

If the King's Library was training data for an AI model instead of an exhibit for public consultation, as an AI developer I'd want to know a few things like where it came from and what happened to it before it got here. Not getting overly obsessed with the linguistics of how well it was written.

The King's Library answers that question better than most archives can because it's construction is unusually well documented. We know which decisions were made and the conditions that enabled it to survive through time. That is what is true.

Most of AI training data doesn't come with anything like that. A scraped archive of books bought in bulk from second-hand booksellers doesn't come with a spreadsheet of the costs or the instructions that were given by the lead developer. Whatever is used is what fell outside of a copyright window, whatever someone with access to hardware and infrastructure could digitise. I've spent most of this piece guessing what's on the shelves. An AI developer training on a scraped archive is doing the same guessing, on an even bigger scale without admitting that they're guessing.

The King's Library and a scraped archive isn't neutral. The difference is that one left a paper trail that could be viewed over 200 hundred years later. The other one doesn't.

Entrance to the British Library, with the building's name above the glass doors

A Note on Searching

I said I don't know what titles are actually in the King's Library but that's easily resolved.

When George IV gifted the collection in 1823, his librarian Frederick Augusta Barnard produced a full five-volume printed catalogue, Bibliothecae Regiae Catalogus (1820–29). It's digitised on the Internet Archive and searchable without a Reader Pass, a list of the whole collection as it stood at George III's death. That's where I'd look first before setting foot in the building.

To then locate and request a physical copy, the British Library's live catalogue can be searched by author. The shelfmark has to be quoted exactly, and need ordering around 48 hours ahead of a Reading Room visit.

Those steps turns "I don't know what's in there" into an answerable question.

Back to blog

Leave a comment

Please note, comments need to be approved before they are published.