Data · dataset · 2026
Chronicling America (US Library of Congress) Historical Newspapers
Listed in Hugging Face Datasets
Description
Chronicling America (US Library of Congress) - Parquet Dataset A high-performance, columnar Apache Parquet dataset containing digitized, OCR-extracted historical American newspapers from the US Library of Congress Chronicling America / National Digital Newspaper Program (NDNP). Produced by streaming and transmuting massive Library of Congress preservation archives (.tar.bz2, METS/MODS, and ALTO XML) into compact, issue-level Parquet shards partitioned by state, newspaper… See the full description on the dataset page: huggingface.co/datasets/Tim-Pinecone/LOC-Chronicling-America.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/Tim-Pinecone/LOC-Chronicling-America ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/Tim-Pinecone/LOC-Chronicling-America ↗
metadata API · from Hugging Face
Topics
- Stated by source
- fill mask · question answering · text retrieval · token classification
- From keywords
- Computer Science & AI · Optical character recognition
Provenance · 1 source records, 13 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | Tim-Pinecone/LOC-Chronicling-America | 7 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[method].local:method:ocr | mapping · Hugging Face | vocabulary-mapper@1.0.0 | keywords['ocr'] |
| concepts[task].hf_task:fill-mask | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:question-answering | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:text-retrieval | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:token-classification | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |