
Researchers often assume that the OCR text displayed by a newspaper database represents the complete article. In reality, many databases index only a portion of the text or create excerpts that end abruptly when a column continues elsewhere on the page. Important details—including names, places, and relationships—may appear just outside the indexed section, causing researchers to overlook information that is actually present in the original newspaper.
Distortion Type
Truncated or Cropped OCR Excerpts
Why it Happens
Databases index snippets, not full columns.
Why it Matters
Names at column edges vanish; partial entries mis-sort under neighboring text.
Search Strategies
- Open the image view to confirm full context.
- Expand date range—an event may appear again.
- If OCR ends abruptly, locate the continuation page.
Example
A researcher discovers an OCR excerpt mentioning a probate notice for Samuel Carter but finds no family information in the indexed text. After opening the full newspaper image, they realize the article continues into the next column, where the names of Samuel’s widow, children, and residence are listed. Because the database captured only part of the article, the most valuable genealogical details never appeared in the searchable OCR results. Viewing the complete page reveals information that would otherwise have been missed.
Key Takeaway
Never assume the OCR excerpt tells the whole story. The most important details are often waiting just beyond the visible snippet.
The Newspaper Distortions Field Guides
Discover all 25 Newspaper Distortions, companion Field Notes, videos, and audio lessons in the Field Guides Collection