Hugging Face Datasets2026 · Image
rustensai/russian-handwriting-ocr…Предназначен для дообучения vision-language моделей (например, Qwen3 VL) на задачу OCR русского рукописного текста.… …1305 Уникальных текстов: 575 Средняя длина текста: 3790 символов Типы изображений… See the full description on the dataset page: https://huggingface.co/datasets/rustensai/russian-handwriting-ocr…
Hugging Face Datasets2026 · Image · gated
5CD-AI/Viet-Handwriting-OCR-v2WE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2024 · Image
Thai Handwriting Dataset…Thai Handwriting Dataset This dataset combines two major Thai handwriting datasets: BEST 2019 Thai Handwriting Recognition dataset (train-0000.parquet) Thai Handwritten Free Dataset…
Hugging Face Datasets2026 · Image · gated
Handwritten GCSE Exam Answers Dataset…Usage Examples Basic Dataset Loading from datasets import load_dataset # Load the full dataset dataset = load_dataset("JunaidMB/handwriting-ocr-images-dataset") # Access specific splits… …train_data = dataset['train'] # 62 samples test_data =… See the full description on the dataset page: https://huggingface.co/datasets/JunaidMB/handwriting-ocr-images-dataset.…
Hugging Face Datasets2026 · Image · gated
5CD-AI/VietHTR-LineWE’RE PREPARING A MORE COMPLETE VERSION AND GETTING THE PAPER READY FOR PUBLICATION...
Hugging Face Datasets2025 · Image
Boston Public Library Card CatalogBoston Public Library Rare Books Card Catalog Dataset Dataset Description This dataset contains approximately 410,000 digitized catalog cards from the Boston Public Library's Rare Books Department card catalog. The cards represent the main entry catalog (author/title cross-referenced) covering printed materials from various historical periods. Why This Dataset? Historical card catalogs are rich so
Hugging Face Datasets2026 · Image
Burmese Handwritten Sentence Dataset (BHSD)…Burmese Handwritten Sentence Dataset (BHSD) BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR… …The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers. This dataset would not have been possible without its volunteers.…
Hugging Face Datasets2022 · Table · Parquet
Berlin State Library OCR…At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.… …For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012).…
Hugging Face Datasets2026 · Image
sarvamai/indic-ocr-bench…Sarvam Indic OCR Bench Global benchmarks focus heavily on English document parsing, and to the best of our knowledge there is no Indic OCR benchmark of comparable breadth and rigor.… …The benchmark covers 23… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-ocr-bench.…
Hugging Face Datasets2026 · Table · Parquet
PubMed-OCR…PubMed-OCR: PMC Open Access OCR Annotations PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs.… …on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.…
Hugging Face Datasets2026 · Table · CSV
Sintético de citações jurídicas (Desafio Jusbrasil BRACIS 2026)Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026 Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso. São duas versões, com o mesmo gabarito (mes
Hugging Face Datasets2025 · Image
MyanmarOCR-ImageText…🇲🇲 MyanmarOCR-ImageText Dataset A clean and diverse Burmese Image-to-Text dataset for OCR and multimodal AI research. 📌 Summary Total images: 41,664 Unique Burmese text entries:… …1,139 Styles per text: 32 variations each Resolution: 512 × 512 File types: PNG/JPG images Dataset split: train only Use cases: OCR, I2T (image-to-text), VLM pretrain/fine-tune All…
Hugging Face Datasets2025 · Image
IIIT5KMETA https://github.com/open-mmlab/mmocr/blob/main/dataset_zoo/iiit5k/metafile.yml Name: 'IIIT5K' Paper: Title: Scene Text Recognition using Higher Order Language Priors URL: http://cvit.iiit.ac.in/projects/SceneTextUnderstanding/Home/mishraBMVC12.pdf Venue: BMVC Year: '2012' BibTeX: '@InProceedings{MishraBMVC12, author = "Mishra, A. and Alahari, K. and Jawahar, C.~V.", title = "Scene Text Recogni
Hugging Face Datasets2022 · Image
Europeana NewspapersDataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiativ
Hugging Face Datasets2026 · Image
AdvSpot…AdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation…
Hugging Face Datasets2022 · Image
chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition
Hugging Face Datasets2026 · Image
ParseBenchParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics d
Hugging Face Datasets2026 · Image · gated
Institutional Newspapers: Boston Public Library…Public Library. 1,473,635 public domain newspaper scans, published between 1795 and 1930 83,147,041 individual crops segmented from those scans 16.3 billion o200k_base tokens of VLM OCR…
Hugging Face Datasets2026 · Image
Ledger…Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values.…
Hugging Face Datasets2026 · Image
krotreaksmey/khmer-math-textbook…🇰🇭 Khmer Math Textbook Line-Level OCR Dataset A large-scale, high-resolution dataset of line-level Khmer text and mathematical formulas extracted from official Cambodian Grade 9,…
Hugging Face Datasets2026 · Text
Manga109-s Text Line AnnotationsManga109-s Text Line Annotations High-precision, line-level bounding box and polygon annotations for the Manga109-s Dataset, supporting both full manga pages and speech bubble crops. Furigana is not labeled and is almost entirely excluded from line labels. Includes 8-point oriented polygons for slanted/rotated text lines. The annotation process is documented in METHODOLOGY.md (WIP). Notice: This d
Hugging Face Datasets2026 · Text
VotingBookletsVotingBooklets Dataset Summary VotingBooklets is a large-scale four-language parallel corpus extracted from the complete collection of Swiss federal voting booklets (Abstimmungsbüchlein), covering federal votes from June 1977 to March 2026. It contains aligned paragraph-level segments across German (de), French (fr), Italian (it), and Romansh Grischun (rm), and serves as a resource for low-resourc
Hugging Face Datasets2026 · Image
Indonesian KTP Dataset 24K (Flat & 3D Perspective Augmented)…synthetic dataset designed to advance SOTA (State-of-the-Art) research in Document Information Extraction (DIE), Key Information Extraction (KIE), and Optical Character Recognition (OCR…
Hugging Face Datasets2026 · Image
Synthetic Text Images (English)Synthetic Text Images (English) A synthetic dataset of rendered text images with rich per-sample annotations: the text itself, its rendering attributes, background description, applied post-processing, and a natural-language caption. Each image is generated by compositing English text over a procedurally generated background with random font, color, position, rotation, blur, brightness and noise.
Hugging Face Datasets2026 · Image
DuckerMaster/Thai-Synth-Receipts…large-scale, highly robust synthetic dataset of Thai commercial documents designed specifically for training and evaluating state-of-the-art Document AI and Optical Character Recognition (OCR…