Publication Intelligence articles
Can AI Evaluate a Book Index?
An Oxford history case study shows how AI can check index coverage, headings, and page references—and why a strong score can still require corrections.
AI can evaluate a book index—if its judgments are tied to the source and its publication-readiness rules are explicit.
That distinction matters because evaluating an index is not the same as checking whether its page numbers exist or whether it resembles another index. A subject index makes thousands of claims about a book: that a concept matters, that a heading represents it accurately, that each locator leads to useful treatment, and that the resulting structure helps a reader navigate.
A book-scale test
We applied a source-grounded, AI-assisted evaluation to the supplied published index for William Doyle’s 2002 second edition of The Oxford History of the French Revolution. The public Oxford study covers 425 supplied body-text pages, 1,904 index records, 5,338 atomic locator claims, and 16 cross-references. The current comparison uses a shared source-first benchmark for expected subjects and treatments.
The evaluation examined the index in both directions. Candidate-to-source review tested whether the entries and locators were supported. Source-to-candidate review asked whether important subjects and expected treatments could actually be found. A whole-index pass then examined hierarchy, density, cross-references, mechanics, and reader navigation.
The score is a weighted quality summary, not “percent correct.” The readiness gates identify defects that stronger results elsewhere cannot cancel:
- Unsupported no-fit locators: a page must support the complete heading and subheading, not just mention a related term.
- Broken references: a delivered see or see-also route must have a valid destination.
The assessment itself is reported as valid; that is separate from the index’s readiness. No human release decision is recorded. The method’s verdict is not a publisher’s approval or rejection.
The evaluated original-index file was distributed by IndexerLabs; its relationship to the publisher’s authoritative index needs to remain explicit.¹
A locator also has to do more than contain the right word. As the Garonne example shows, every page reference makes its own promise to the reader. The Oxford audit tests that promise against the complete heading path, not an isolated term.
Checks, judgments, and publication decisions
- Mechanical validation. Software can map pages, expand ranges, validate schemas, account for denominators, resolve identifiers, and reproduce declared arithmetic.
- Source-grounded evaluation. Evidence-grounded review assesses whether passages provide useful treatment, whether headings preserve meaning and stance, and whether important access is missing.
- Method readiness and human release. Separate gates determine readiness under the declared method. A human release decision remains a distinct record; this study does not record one.
The Subject Index Evaluation Standard keeps these layers separate. The score is calculated from frozen ledgers; diagnostic item grades do not add up to it, and publication gates remain outside the score.
Why evaluation can be harder than generation
Generation proposes one possible navigation system. Evaluation must verify every proposed claim and search the book for important access that was never proposed. It has to build and review a source-first benchmark, expand every range into individual page assignments, audit the candidate in both directions, inspect the index as a whole, preserve uncertainty, and validate the calculation.
That structure explains why evaluation can require more computation than generation, especially when several candidates share the same benchmark. The Oxford work did not record comparable token use, hardware, cost, or human minutes for generation and evaluation, however, so it does not establish an empirical cost ratio.
So, can AI evaluate a book index?
Yes. In the Oxford study, AI agents reviewed source evidence at book scale, applied a declared rubric, surfaced contradictions, and produced complete audit ledgers. Deterministic software then validated the ledgers and calculated the score and gate outcomes.
This article follows the supplied published index. The wider Oxford study now compares four candidates against a shared source benchmark, but still concerns one English-language scholarly history. It does not establish reliability across genres or model versions, and live-reader testing measures something the rubric does not.
There is also a direct competing interest: I created the methodology and develop Publication Intelligence. The benchmark was prepared in candidate-unseen contexts, but the study is not institutionally independent. Its defense is therefore transparency: the long-form methodology, evaluation framework, frozen benchmark, public result artifacts, and item-level findings are available for scrutiny.