Newspaper Distortions – The Careless Cutter

Researchers often assume that the OCR text displayed by a newspaper database represents the complete article. In reality, many databases index only a portion of the text or create excerpts that end abruptly when a column continues elsewhere on the page. Important details—including names, places, and relationships—may appear just outside the indexed section, causing researchers to overlook information that is actually present in the original newspaper.

Distortion Type

Truncated or Cropped OCR Excerpts

Why it Happens

Databases index snippets, not full columns.

Why it Matters

Names at column edges vanish; partial entries mis-sort under neighboring text.

Search Strategies

  • Open the image view to confirm full context.
  • Expand date range—an event may appear again.
  • If OCR ends abruptly, locate the continuation page.

Example

A researcher discovers an OCR excerpt mentioning a probate notice for Samuel Carter but finds no family information in the indexed text. After opening the full newspaper image, they realize the article continues into the next column, where the names of Samuel’s widow, children, and residence are listed. Because the database captured only part of the article, the most valuable genealogical details never appeared in the searchable OCR results. Viewing the complete page reveals information that would otherwise have been missed.       

Key Takeaway

Never assume the OCR excerpt tells the whole story. The most important details are often waiting just beyond the visible snippet.

The Newspaper Distortions Field Guides

Discover all 25 Newspaper Distortions, companion Field Notes, videos, and audio lessons in the Field Guides Collection

Leave a Reply

Your email address will not be published. Required fields are marked *