Publication Intelligence articles
How to Compare Human and AI Indexes Fairly
A practical checklist for comparing human and AI book indexes: select exact artifacts, match the source, share the rules and build an independent benchmark.
You have a published human index and an AI-generated index for the same book. One is longer; the other looks cleaner. Which serves readers better? Neither appearance nor agreement with the published index can settle that question. First establish what the comparison can fairly measure.
Before scoring: six checks
- Question: Are you comparing finished indexes, unedited drafts, or the work needed to reach publication?
- Selection: Choose each candidate by a stated rule before inspecting its quality.
- Identity: Preserve the exact source and index files, including their edition and editing history.
- Scope: Confirm which pages and other material were supplied and can be assessed.
- Rules: Use the same audience, indexing policy, rubric and review procedure.
- Benchmark: Prepare a map of important subjects and useful passages from the source without consulting the candidates.
Choose the artifact that answers your question
“AI versus human” is too broad to be a useful study design. A comparison of raw AI output with a professionally edited human index can answer a practical question: how usable is the automated draft beside the published result? It cannot isolate the effect of authorship when the workflows include different amounts of review.
If you want to compare complete production workflows, record the editing required, elapsed time and cost as well as final quality. Keep those measures separate: a strong finished index may have required substantial repair.
When several AI outputs are available, choose prospectively—for example, the artifact attached to a specified public demonstration, or the first completed run under fixed settings. Record why it qualifies. Selecting whichever looks best or worst after inspection changes the question and invites cherry-picking.
Match the edition—and the supplied material
The same title does not establish the same source. Revised editions can change arguments, chapters and pagination. Even within one edition, a PDF page position may differ from the printed page number. Verify that the locators refer to the source you actually have.
Save exact copies of each candidate and source, the retrieval date, their origin and any disclosed corrections. A file checksum helps establish that later reviewers are examining the same bytes; it does not prove that the file is the publisher’s authoritative original. Check the rendered index against extracted text before attributing a damaged accent or lost indentation to its creator.
Record missing pages, appendixes, notes, captions or illustrations. If the supplied source ends before a locator’s destination, mark that claim outside the assessable scope rather than silently grading it wrong. If one candidate received less material, disclose that difference and use a common scope where defensible. A body-text-only audit supports a body-text conclusion.
Let the book define the benchmark
A published human index is an editorial artifact with its own choices, constraints and possible errors. Treat it as another candidate. Using its entries as the answer key automatically favors its selection and wording, while hiding subjects that both indexes omit.
Instead, read the source without the candidate indexes. Record important concepts, their relative significance, useful treatments, relationships and likely reader questions. Allow defensible alternative terms and structures. Freeze this benchmark before scoring. If later evidence requires a substantive change, document the revision and reassess every candidate affected by it.
If both indexes omit the discussion of tariffs, agreement between them will not reveal the omission. The independently prepared source map can. Conversely, adding page 75 to the metering-effects entry needs support for that complete heading: proximity to the discussion is insufficient.
Preparation without seeing candidates protects against copying their choices. It does not establish institutional independence. Disclose who designed the method, who reviewed the benchmark and any competing interest.
Apply one set of rules in both directions
Agree on intended readers, length constraints, heading depth, treatment of notes and names, locator conventions, error severity and scoring weights. The American Society for Indexing’s evaluation checklist offers useful questions about reader appropriateness, headings, locators, cross-references and presentation. Convert the relevant questions into declared rules for this comparison.
Then examine each candidate in both directions. Candidate-to-source review asks whether its headings and page references are justified. Source-to-candidate review asks whether important subjects and treatments can be found. A tidy index can fail through omission; a large index can fail through clutter. Entry counts alone measure neither.
Use neutral candidate labels where practical and the same review procedure for each. Record disagreements and resolve them against the source. If you sample, disclose the selection method and limits; checking a few plausible entries cannot establish whole-index publication readiness.
Write the conclusion your evidence supports
Report consequential omissions, misleading headings, unsupported locators, navigation problems and repair needs alongside any total score. Name the exact artifacts and shared benchmark. State exclusions and unresolved judgments. A result for one book and one run does not establish a universal ranking of human indexers or AI systems.
Before accepting a comparison, check whether another reviewer could recover the same materials and understand every rule. If candidate identity, source coverage or benchmark independence is unclear, settle that uncertainty before using the ranking to make an editorial decision. For a fuller account of this approach, see our source-grounded evaluation methodology.