The U.S. National Archives holds more than a billion electronic files, in over 700 distinct file formats, totaling more than a petabyte of records that someone, at some point, decided were worth keeping forever. That volume did not accumulate gradually. It is the direct consequence of a world that now generates its historical record natively in digital form, faster than any archive staffed at human scale can process, describe, or make searchable.
The displacement risk for archivists is classified as Moderate, with a horizon toward 2042. That reflects a genuine split in the job. AI is already doing real work on one half of it, reading, transcribing, and tagging the growing backlog. It has made almost no progress on the other half: deciding which records, out of everything a government, an institution, or a family produces, actually deserve to survive.
That distinction, between processing a record and judging its value, is where this profession's future will actually be decided. Automating the first has already started at scale. The second is the part every major archival body, from the Society of American Archivists to the International Council on Archives, is still explicitly unwilling to hand to a model.
Key Points
- Archivists are rated Moderate risk with a displacement horizon toward 2042, split between transcription and metadata work that AI already automates well and appraisal judgment that it does not.
- Transkribus, the leading handwritten text recognition platform, is used by more than 2,000 archives and libraries and has processed over 200 million pages, with trained models reaching above 95% accuracy on clean documents.
- Accuracy drops sharply on the hardest material: published error-rate studies show 10 to 15% character error rates on damaged, unusually handwritten, or historically obscure documents, the exact material archivists are most often asked to rescue.
- Archival appraisal, deciding what to keep out of an overwhelming volume of records, remains a human judgment task; researchers and professional bodies warn that training AI on past appraisal decisions risks automating historical bias rather than removing it.
- The U.S. Bureau of Labor Statistics projects 4% employment growth for archivists, curators, and museum workers through 2035, driven explicitly by organizations needing to manage expanding digital records, not by automation reducing headcount.
What an Archivist Actually Does
The job covers the full lifecycle of a record after it stops being actively used: appraising which materials have enduring historical, legal, or evidentiary value, arranging and describing collections so researchers can find what they need, preserving fragile physical and digital formats against decay and obsolescence, and building the finding aids and catalog structures that make a collection navigable rather than just stored. Increasingly, it also means managing records that were never physical at all, email, databases, government systems, born-digital documents that arrive in formats no card catalog was ever designed for.
What AI Is Already Doing
Transkribus, developed by the READ-COOP cooperative, is the clearest success story in the field. The platform is in use at more than 2,000 archives and libraries, has processed over 200 million pages, and supports more than 300 trained models across historical scripts and languages. Institution-specific trained models routinely exceed 95% accuracy on well-preserved material. The U.S. National Archives runs OCR through open-source Tesseract software as records enter cloud storage, automating basic searchability at ingestion rather than as a separate project. The Library of Congress runs a hybrid model through its By the People and FixIt+ programs: AI generates a first-pass transcript, and volunteers correct and refine it, treating automation as a starting point rather than a finished product.
THE BACKLOG
NARA's own strategic planning documents cite more than a billion electronic files, across 700-plus file formats, as the scale of what the agency is responsible for preserving. No archival staff, at any plausible funding level, processes that volume by hand. AI is not displacing archivists from a stable job. It is being deployed against a backlog that already exceeded human capacity before generative AI existed.
Metadata and discovery are moving the same direction. Europeana's AI4Culture platform supports metadata enrichment across the 26 million objects the European digital library aggregates from more than 2,200 institutions. Semantic search systems, built on embeddings rather than keyword matching, increasingly let researchers query archives by meaning, find everything related to a historical event or theme, rather than by the exact words a decades-old catalog entry happened to use.
Where AI Still Struggles
The accuracy numbers that make headlines describe the easy cases. Published research on handwritten text recognition shows character error rates of 3 to 5% on well-preserved, regular historical handwriting, the kind Transkribus was largely trained on. That number climbs to 10 to 15% on damaged parchment, unusual hands, heavy abbreviation, or scripts the model has seen little of, and one widely cited public model for Italian administrative handwriting from 1550 to 1700 posts a 9.15% error rate even on its own benchmark. A 15% error rate sounds tolerable until it is applied to a legal document, a land deed, or a name in a genealogical record, where a single misread character changes the meaning of the page.
The deeper problem is coverage, not just accuracy. Large language models are trained overwhelmingly on modern text. Historical language, deprecated spelling conventions, regional dialects, and scripts with thin surviving corpora produce exactly the low-accuracy, low-confidence output that requires an archivist's trained eye to catch, which is why institutions from the Library of Congress to national archives across Europe have settled on hybrid human-AI transcription rather than fully automated pipelines.
The Judgment Call AI Can't Make
Appraisal, the decision about what to keep and what to discard, is where the profession's ethical stakes concentrate. Research on AI-assisted appraisal describes an emerging consensus that automation can help surface records of likely significance inside enormous email archives or government record systems, but that the actual decision requires human discretion that current models cannot substitute for. The risk runs in a specific direction: a model trained on what past archivists chose to preserve will reproduce whatever historical bias shaped those choices, automating exclusion rather than correcting it, unless a person is actively auditing the outcome.
Professional bodies have been explicit about where that line sits. The Society of American Archivists' Code of Ethics, revised as recently as 2025, centers professional judgment, authenticity, and accountability as core obligations that cannot be outsourced. The International Council on Archives has gone further, arguing in a submission to the World Intellectual Property Organization that AI-related copyright and access restrictions could actively damage archives' ability to preserve public accountability and historical transparency, positioning the profession as a check on AI's downstream effects on public memory, not merely a consumer of AI tools.
The Labor Market Reality
The U.S. Bureau of Labor Statistics projects 4% employment growth for archivists, curators, and museum workers through 2035, roughly in line with the average occupation, with about 4,300 openings a year and a median salary of $60,330. The agency's own stated driver for that growth is telling: organizations need more people to manage expanding digital records, not fewer. That is the opposite of the pattern AI Doomsday has documented in professions where AI reduces the total volume of necessary human labor. Here, the volume of material needing archival attention is growing faster than any single technology has managed to compress it.
How to Use AI as an Archivist Now
For transcription: treat platforms like Transkribus as a first pass on volume, not a final answer on accuracy. The material most worth an archivist's direct attention is exactly the material where error rates spike, damaged, unusual, or historically obscure documents.
For metadata and discovery: semantic search and embedding-based tools are genuinely useful for surfacing connections across large, previously unindexed collections. Use them to find what a keyword search would miss, then verify the result against the source.
For appraisal: use AI to flag records that pattern-match as significant inside an overwhelming volume, an email archive, a records transfer, a born-digital government system, but treat every flag as a starting point for review, not a decision. Document your reasoning where the model's training data might carry a bias you would not otherwise catch.
What I Think
The 2042 horizon looks reasonable to me because the two halves of this job are moving at genuinely different speeds. Transcription and metadata work will keep getting faster and cheaper, and archivists who resist that shift are going to lose ground to institutions that don't. Appraisal is a different problem entirely, because it is not really a technical task being performed slowly. It is a judgment about historical value that reflects the judging institution's priorities, and no model improvement changes who is accountable for that judgment.
What I find most compelling about this profession's position is that its own professional bodies got ahead of the risk rather than reacting to it. The concern about training data encoding old biases into what gets preserved is not a hypothetical raised by outside critics. It is coming from archivists themselves, which suggests the field understands its own exposure better than most. The backlog is real, the tools are useful, and the profession's insistence on keeping a human in the loop for the decisions that matter most is, so far, holding.
"AI can read a billion pages faster than any archivist ever could. It still can't tell you which of them matter in fifty years. That judgment is the whole job, and no one has figured out how to automate accountability for it."