{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:visual-question-answering",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Image
OmniTaskonomy Recipe DataOmniTaskonomy Recipe Data Paired image-to-image (I2I) and image-to-text (I2T) tasks for the R1–R6 training recipes and gradient analysis in OmniTaskonomy. Each of the six subsets has train and val splits. One row contains both objectives for the same task instance. from datasets import load_dataset data = load_dataset("Wakals/OmniTaskonomy_Recipe_Data", "jigsaw", split="train", streaming=True) sam
Hugging Face Datasets2026 · Image
OmniTaskonomyOmniTaskonomy OmniTaskonomy groups visual tasks into Recognition, Reconstruction, and Reorganization. Each task and modality has its own split, named family__task__i2i or family__task__i2t. The i2t and i2i configs group the evaluation and training splits, respectively. This release contains 9,444 I2T evaluation samples across 25 tasks and 350,000 I2I training samples across 7 tasks. I2T rows have
Hugging Face Datasets2026 · Image
VietTravelVQA v2VietTravelVQA v2 VietTravelVQA v2 is a Vietnamese visual question answering dataset about tourism and cultural heritage in Vietnam. This release contains 9,530 question-answer pairs associated with 1,406 images. It combines the original 7,030 annotated pairs with 2,500 additional knowledge-grounded pairs. Dataset summary Split Question-answer pairs Images Train 6,805 1,051 Validation 1,010 194 Tes
Hugging Face Datasets2026 · Text · gated
PathoVernierPathoVernier PathoVernier is a benchmark for quantitative cell-composition reasoning on H&E histopathology patches. Each question requires counting specific nucleus types in specific image regions and deriving an answer from those counts with a deterministic rule. Reference counts come from expert nucleus annotations, so a model's reported counts can be checked in addition to its final answer. 759
Hugging Face Datasets2026 · Image
PhysAlignPhysAlign PhysAlign is a bilingual multimodal physics benchmark for testing whether a model can make local observations and bind textual or visual references to the correct candidate entity. It does not ask the evaluated model to solve the original examination problem. PhysAlign was constructed from selected examples in previously released public benchmarks, followed by source adaptation, probe ge
Hugging Face Datasets2026 · Text
LuckerZ/MoGroundMoGround Three resources for measuring modality distraction: a model answers a question correctly from one modality alone, then answers it wrongly once answer-irrelevant context arrives in the other modality. Every item is checked to be answerable from exactly one modality, so the distraction it measures cannot be explained by the question being unanswerable. pool items what it is moground_base 34
Hugging Face Datasets2026 · Image
wakinghours/PanoVQAPanoVQA PanoVQA provides English visual question-answering annotations for panoramic scenes from NuScenes, DeepAccident, and BlendPASS. Questions cover topics such as scene descriptions, object attributes, spatial relationships, visibility, and traffic situations. The full dataset contains 653,914 question-answer records across 40,910 images. A smaller version, PanoVQA_mini, contains 13,427 record
Hugging Face Datasets2026 · Image
ImajevBench v2.0-lite (preview)ImajevBench v2.0-lite (preview) A benchmark for typed decisions made from a photo, a written rule, or both. A system receives the evidence (0–2 images and a state: an order, a rule, a form, a claim) and one typed question, and must return exactly one of: yes/no, a listed choice, an integer level, or Unknown when the evidence does not determine the answer. It checks four things: whether the decisio
Hugging Face Datasets2026 · Image
krotreaksmey/khmer-math-textbook🇰🇭 Khmer Math Textbook Line-Level OCR Dataset A large-scale, high-resolution dataset of line-level Khmer text and mathematical formulas extracted from official Cambodian Grade 9, 10, 11, and 12 mathematics textbooks. 📊 Dataset Summary Total Samples: 29,027 labeled line crops Coverage: Grade 9 (New): 6,847 line crops (math-G9-1 to math-G9-6847) Grades 10, 11, 12: 22,180 line crops Columns: Strictly
Hugging Face Datasets2026 · Text
TRACE v1.1.0TRACE Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding Timestamped video QA and proactive-response annotations for evaluating what a model knows, when it knows it, and how it responds. How TRACE works TRACE separates causal video delivery, model interaction, and scoring. The same public contract makes QA and Proactive Response results auditable across models: QA: answ
Hugging Face Datasets2026 · Text · gated
WearerTextBenchWearerTextBench WearerText: Benchmarking Wearer-Centered Scene Text Understanding in AI-Glasses Videos Project & code · Project page Release scope This repository contains the test split only: 107 videos and 1,391 QA pairs, with 107 examples for each of 13 tasks. Each video is stored once and referenced by 13 metadata rows. Videos occupy 2,471,387,363 bytes. The original files use MPEG-4 Part 2 vi
Hugging Face Datasets2026 · Image
DAREBench: Deployment-Aware and Reliable Evaluation of Models as AgentsDAREBench DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents) is a workload- and deployment-aware benchmark for evaluating models as agents. Built on a shared OpenClaw execution environment, it organizes 233 tasks selected and adapted from 22 source benchmarks into a 2×3 workload matrix defined by input modality and execution form, and evaluates them under a unified contract-b
Hugging Face Datasets2026 · Image
AdvSpotAdvSpot AdvSpot is the first grounded adversarial OCR benchmark, introduced in the paper ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation. It targets visual text patterns that are readable to humans but challenging for multimodal large language models (MLLMs) to localize and recognize, and evaluates models with region-level annotations: bounding boxes,
Hugging Face Datasets2026 · Image
TraceSpatial-TraceTraceSpatial-Trace Project · Paper · Code TraceSpatial-Trace is the RGB-and-QA tracing subset of TraceSpatial, introduced by the RoboTracer project. It supports learning to translate language instructions into spatial waypoints for object manipulation and robot end-effector motion. This release contains 517,215 referenced RGB images and 3,623,880 question–answer pairs, extracted from 1,067,822 ori
Hugging Face Datasets2026 · Table · Parquet
VLADBench-reevalVLADBench-reeval We re-evaluated the VLADBench benchmark (Li et al., 2025, arXiv:2503.21505) against current SOTA VLMs under the original scoring criteria and prompts. See Eventual-Inc/VLADBench for the code, the companion site for interactive results, and the article for a summary of the findings. Cost vs Score Up and to the left is better. The dashed line is the cost-performance frontier which i
Hugging Face Datasets2026 · Image
AgroOmniAgroOmni AgroOmni is a large-scale multi-view agricultural dataset introduced in "AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning" (arXiv:2603.14342). It spans ground, UAV, and satellite imagery to address the ground-level bias of existing agricultural multimodal models, with 288K visual question answering pairs covering 56 specialized task categories a
Hugging Face Datasets2026 · Image
LSI-108KLSI-108K Interaction-derived supervision for spatial reasoning LSI-108K is the dataset introduced in Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World. It contains 107,518 verifiable QA pairs organized as a three-level spatial interaction curriculum. Overview Presentation A three-level curriculum Split Samples Curriculum role L1 15,109 Passive… S
Hugging Face Datasets2026 · Image
Metric-Bench TestMetric-Bench Test Metric-Bench Test set evaluates metric spatial understanding from indoor RGB images and explicit anchor measurements. Each example asks for a physical measurement of a referred object or the distance between two objects. Inputs consist of an image and an English question; the reference answer is a numeric JSON object. This release contains 1,340 questions, 134 images and 20 scene
Hugging Face Datasets2026 · Image
RekaAI/RekaDaily-10k-processedRekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second per
Hugging Face Datasets2026 · Image
PDFA OCR Dataset - KREATIVE TIME BOXPDFA OCR Dataset Curated and Published by KREATIVE TIME BOX This dataset contains document page images along with their corresponding OCR layout bounding box annotations derived from PDFA document extraction pipelines. Dataset Overview Organization / Creator: KREATIVE TIME BOX Images: 27,499 PNG files (~8.8 GB) JSON Annotations: 6,989 JSON files (~52 MB) Image Format: PNG (RGB document page render
Hugging Face Datasets2026 · dataset
LogiScope-VQA (Preview)Dataset Card for LogiScope-VQA (Preview) ⚠️ This is a preview release. It contains approximately 10% of the full LogiScope-VQA benchmark, provided for academic review during the submission period. The complete version (10,274 VQA pairs) will be publicly released after paper acceptance. Dataset Description LogiScope-VQA is the first dedicated benchmark for evaluating Large Multimodal Models (LMMs)
Hugging Face Datasets2026 · Image · gated
taesiri/BlindLoop-GenerationsBlindLoop Generations A flat, general-purpose visual-question-answering dataset generated by the BlindLoop paper experiments. Each row is one concrete question instance with a native Hugging Face Image value, question, gold answer, answer choices, and fully filterable generation provenance. Config Rows Tasks Unique source images section1_all 516,810 1,301 249,488 section2_all 335,271 875 166,364 c
Hugging Face Datasets2026 · dataset · gated
nuReasoningnuReasoning Paper Website nuReasoning Website nuReasoning Dev-kit nuReasoning is a reasoning-centric multimodal autonomous driving dataset for evaluating and training end-to-end driving systems in long-tail real-world scenarios. Each sample is built around a driving clip with synchronized multi-camera images, LiDAR data, ego state, object annotations, HD map, routing, and frame-level reasoning ann
Hugging Face Datasets2026 · Image
Biomedica 2025 Eval SetBiomedica 2025 Eval Set Unified test snapshot of the BioMedica 2025 vision–language evaluation suites used in AMInZeroShotOpenEvalAllTasks. Every row is a single image with closed-ended options, the gold answer, and provenance fields that point back to the original dataset. Images are stored as original JPEG/PNG bytes (or JPEG-encoded arrays) inside parquet so the Hugging Face dataset viewer is en
Hugging Face Datasets2026 · Table · Parquet
datalab-to/omni_extract_benchOmni Extract Bench We weren’t satisfied with the current benchmarking options for extraction. They were biased, didn’t use realistic data and were hard to audit. Our view is that an extraction benchmark should do two things: Help customers choose the right vendor; and Give engineers a way to diagnose what’s actually going wrong in a given model. That’s why we built OmniExtractBench. OmniExtractBen