Hugging Face Datasets2026 · Image
DuckerMaster/Thai-Synth-ReceiptsThai-Synth-Receipts Thai-Synth-Receipts is a large-scale, highly robust synthetic dataset of Thai commercial documents designed specifically for training and evaluating state-of-the-art Document AI and Optical Character Recognition (OCR) models. The dataset consists of 14,976 high-resolution document images (Receipts, Tax Invoices, Thermal Slips, and Quotations) across three distinct degradation v
Hugging Face Datasets2026 · Image
VidScribeVidScribe VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12). Release status. This repository hosts the public hal
Hugging Face Datasets2026 · Image
krotreaksmey/khmer-math-textbook🇰🇭 Khmer Math Textbook Line-Level OCR Dataset A large-scale, high-resolution dataset of line-level Khmer text and mathematical formulas extracted from official Cambodian Grade 9, 10, 11, and 12 mathematics textbooks. 📊 Dataset Summary Total Samples: 29,027 labeled line crops Coverage: Grade 9 (New): 6,847 line crops (math-G9-1 to math-G9-6847) Grades 10, 11, 12: 22,180 line crops Columns: Strictly
Hugging Face Datasets2026 · Table · CSV
Sintético de citações jurídicas (Desafio Jusbrasil BRACIS 2026)Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026 Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso. São duas versões, com o mesmo gabarito (mes
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.
figshare + Loughborough Research Repository2026 · Astronomical catalogue
LongHisDoc: A Comprehensive Benchmark for Chinese Long Historical Document Understanding<p dir="ltr">In this repo, we present <b>LongHisDoc</b>, a pioneering benchmark specifically designed to evaluate the capabilities of LLMs and LVLMs in long-context historical document understanding tasks. This benchmark includes 101 historical documents across 10 categories, with 1,012 expert-annotated question-answer pairs covering four types, and the evidence for the questions is drawn from thr
Hugging Face Datasets2026 · dataset
Chronicling America (US Library of Congress) Historical NewspapersChronicling America (US Library of Congress) - Parquet Dataset A high-performance, columnar Apache Parquet dataset containing digitized, OCR-extracted historical American newspapers from the US Library of Congress Chronicling America / National Digital Newspaper Program (NDNP). Produced by streaming and transmuting massive Library of Congress preservation archives (.tar.bz2, METS/MODS, and ALTO XM
Hugging Face Datasets2026 · Image
AdvSpotAdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation. It targets visual text patterns that are readable to humans but challenging for multimodal large language models (MLLMs) to localize and recognize, and evaluates models with region-level annotations: bounding boxes,
e-cienciaDatos2026 · dataset · unknown
Minitutorial de OCR con un VLM formato taller: fine-tuning, inferencia y métricas con Gemma 3 4B (ipynb + instrucciones)El presente recurso se trata de un minitutorial en formato taller que pretende ser una introducción a realizar OCR con un VLM (Vision Language Model) empleando Python, el repositorio de Hugging Face y Unsloth. Se trata de un cuaderno de Jupyter en formato ipynb (interactive Python notebook) acompañado de un informe técnico que sirve como acompañamiento a este, aunque el ipynb es suficiente. Archiv
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOXPDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from PDFA document extraction pipelines. Dataset Overview Organization / Creator: KREATIVE TIME BOX Images: 27,499 PNG files (~8.8 GB) JSON Annotations: 6,989 JSON files (~52 MB) Image Format: PNG (RGB document page render
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR), error analysis, robustness testing, and research on low-resource scripts. The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)Indonesian KTP Dataset 24K (Flat & Augmented - Commercially Safe) Welcome to the Indonesian KTP (Kartu Tanda Penduduk) Dataset. This is a highly robust, high-fidelity, and commercially safe synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR) specifically for Indone
Hugging Face Datasets2026 · Image · gated
5CD-AI/VietHTR-LineWE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2026 · Image · gated
Institutional Newspapers: Boston Public Library📰 Institutional Newspapers: Boston Public Library A structured dataset derived from the Boston Public Library's public domain newspapers collection, produced by the Institutional Data Initiative in collaboration with Boston Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens o
Hugging Face Datasets2026 · Table · Parquet
artefactory/ledger-long-context-KPI-QALEDGER — Long-Context KPI Question Answering & Page Retrieval This dataset is part of the LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval) benchmark. It supports two of the three LEDGER tasks: Page-level KPI retrieval — given a natural-language question about a financial KPI and the corresponding annual report, retrieve the relevant page(s). Each row includes TRE
Hugging Face Datasets2026 · Image
Ledgerthe LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-
Hugging Face Datasets2026 · Image
sarvamai/indic-ocr-benchSarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor. Sarvam Indic OCR Bench fills this gap with 6,909 curated text-block samples drawn from document pages spanning the 19th century to the present, across a wide range of scan quality and content types—including textbooks,
Hugging Face Datasets2026 · Text
Manga109-s Text Line AnnotationsManga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This d
Hugging Face Datasets2026 · Image · gated
5CD-AI/Viet-Handwriting-OCR-v2WE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2026 · Image
PureDocBenchMain Leaderboard 58 models · 3 matched tracks · 🏆 Search, filter & sort the leaderboard → The current evaluation covers 13 pipeline / multi-stage specialists, 19 end-to-end specialists, and 26 general-purpose VLMs. Each track contains 1,475 pages. The top 10 by the three-track mean, Avg₃, are shown below. Rank Model (release) Type Clean ↑ Digital ↑ Real ↑ Avg₃ ↑ 1 GLM-5.3-Flash (2026-08) General V
Hugging Face Datasets2026 · Text
VotingBookletsVotingBooklets Dataset Summary VotingBooklets is a large-scale four-language parallel corpus extracted from the complete collection of Swiss federal voting booklets (Abstimmungsbüchlein), covering federal votes from June 1977 to March 2026. It contains aligned paragraph-level segments across German (de), French (fr), Italian (it), and Romansh Grischun (rm), and serves as a resource for low-resourc
Hugging Face Datasets2026 · Image
ParseBenchParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics d
Hugging Face Datasets2026 · Image · gated
OCRGenBenchOCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities 🔐 Dataset Access This dataset is gated. To download OCRGenBench, please submit an access request: 👉 Apply for Access — click the Access button on the dataset page Applications are automatically approved. You will receive an email confirmation once access is granted. 📖 Overview OCRGenBench is the most comprehensive be
Hugging Face Datasets2026 · dataset
OCR Synthetic Multilingual v1OCR-Synthetic-Multilingual-v1 Dataset Description Large-scale synthetically generated OCR training dataset for multilingual text detection and recognition. The data was produced using a heavily modified and extended version of SynthDoG (Synthetic Document Generator), originally introduced in the Donut project by Kim et al. This dataset was used to train Nemotron OCR v2, a state-of-the-art multilin