{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:question-answering",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Table · Parquet
OpenJevData-140kDataset Card for OpenJevData-140k Dataset Summary OpenJevData-140k is a curated release of the data collection used to train OpenJev-4B. It contains 146,738 decision-making examples across 19 task categories, organized into SFT and RL splits. Each example presents a state, a question, and a request-specific set of natural-language options. The data include hard answers and soft probability distrib
Hugging Face Datasets2026 · Text
GSM8K-UZ (Cyrillic)GSM8K-UZ (Cyrillic) Uzbek GSM8K in the Cyrillic script. Part of a parallel pair: kurbanovxurshidbek/gsm8k-uz-lat and kurbanovxurshidbek/gsm8k-uz-cyr. Both contain exactly the same problems (same idx), differing only in script. Split Rows train 7417 test 1308 Source and construction Latin text: NeuronUz/gsm8k-uz, a machine translation of openai/gsm8k (main) into Uzbek Latin. Cyrillic text: automati
Hugging Face Datasets2026 · Text
HABIT-BenchHABIT-Bench Evaluating Habit Induction from Longitudinal Weak Evidence in Agent Memory HABIT-Bench evaluates whether a memory system can infer a latent habit from repeated, individually incomplete observations and apply it only when the current context supports it. Tasks test weak-evidence induction, applicability boundaries, local exceptions, temporal changes, and the distinction between user-end
Hugging Face Datasets2026 · Text
The First And Best Fable 5.1 Reasoning DataDataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces
Hugging Face Datasets2026 · Text
GSM8K-UZ (Latin)GSM8K-UZ (Latin) Uzbek GSM8K in the Latin script. Part of a parallel pair: kurbanovxurshidbek/gsm8k-uz-lat and kurbanovxurshidbek/gsm8k-uz-cyr. Both contain exactly the same problems (same idx), differing only in script. Split Rows train 7417 test 1308 Source and construction Latin text: NeuronUz/gsm8k-uz, a machine translation of openai/gsm8k (main) into Uzbek Latin. This dataset reproduces the L
Hugging Face Datasets2026 · Table · Parquet
Qwen3.8-Max DistillationQwen3.8-Max Distillation A quality-filtered derivative of Qwen3.8-Max Distillation 50K by r0b0tlab, prepared for local training and fine-tuning on consumer hardware. This repository takes the original 49,772-example dataset and produces a substantially smaller training set focused on coding, reasoning, instruction following, and tool use. [!CAUTION] Terms and provenance notice — not cleared for un
Hugging Face Datasets2026 · Text
BaRe-Mem DataBaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation Overview In multi-agent systems, a central model can consult advisors, but advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. BaRe-Mem is an online Bayesian reliability memory for multi-agent consultation: it estimates each advisor's reliability fr
Hugging Face Datasets2026 · Image
VietTravelVQA v2VietTravelVQA v2 VietTravelVQA v2 is a Vietnamese visual question answering dataset about tourism and cultural heritage in Vietnam. This release contains 9,530 question-answer pairs associated with 1,406 images. It combines the original 7,030 annotated pairs with 2,500 additional knowledge-grounded pairs. Dataset summary Split Question-answer pairs Images Train 6,805 1,051 Validation 1,010 194 Tes
Hugging Face Datasets2026 · Text
HUMMBL 40k Multi-Agent Wicked Problems & Coordination CorpusHUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet. Dataset Overview The dataset provides structured visibility into how autonomous agents navigate comple
Hugging Face Datasets2026 · Table · Parquet
SecondState FAB — Agent Traces and GradingFAB — Agent Traces and Grading 600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included. The benchmark dataset contains the shared data
Hugging Face Datasets2026 · Table · Parquet
MMDMDataset Card for MMDM Dataset Description Dataset Summary MMDM (Massive Multitask Decision Making) evaluates decisions over request-specific, natural-language options. It contains 4,961 evaluation examples across 17 broad task categories, combining source-grounded questions with constructed decision problems. Candidate sets contain 2–77 options. Option names and criteria are part of the input; the
Hugging Face Datasets2026 · Text
SecondState FAB — Finance Agents BenchmarkFAB — Finance Agents Benchmark FAB is an open-source project for benchmarking LLM agents' ability to perform financial due diligence in a synthetic company data room. FAB consists of a dataset of tasks containing agent instructions, documents and rubrics, and an execution harness for running and evaluating agents. This repository contains the dataset; the harness is available on GitHub. Dataset 50
Hugging Face Datasets2026 · Table · Parquet
MINTQAMINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge MINTQA evaluates how well LLMs answer complex multi-hop (1–4 hops) questions that involve unpopular and new knowledge. It has two subsets: Config Examples Dimension Per-hop type labels pop (MINTQA-POP) 17,887 Knowledge popularity: unpopular vs. popular facts R, F ti (MINTQA-TI) 10,479 Time: new facts (on
Hugging Face Datasets2026 · dataset
Doctor-Patient Conversations — All Human Diseases (Opus 5.5)Opus-5.5 generated Doctor-Patient Conversations for All Human Diseases The sequel to nisten/opus-doctor-patient-conversations-all-human-diseases (Opus 4.8). Same disease list, same 20-key schema, same ChatML conversations — regenerated from scratch with Claude Opus 5.5, one agent per disease, and held to a much stricter bar. Covers every human disease listed on my previous work here: nisten/all-hu
Hugging Face Datasets2026 · Table · CSV
Saeid Homayoun Accounting and Audit AI PortfolioSaeid Homayoun — Accounting & Audit AI Portfolio A curated starting point for public work in AI-enabled accounting, auditing, assurance, finance, SEC/XBRL, ICFR, CAM/KAM, ESG and governance. Start here NAAIL OpenLab SEC 10-Company Accounting Panel Kimi + DeepSeek Accounting/Finance/Audit Stack SEC CAM/ICFR/Governance/ESG Frankenstein Accounting/Audit Benchmark The machine-readable portfolio.csv ra
Hugging Face Datasets2026 · Text
H2S-Research/Highlight-Then-SummarizeHighlight-Then-Summarize Highlight-Then-Summarize (H2S) is a compress-then-reason approach for long-context understanding. It makes evidence localization and information integration explicit before final-answer generation: long document + question -> evidence -> question-conditioned summary -> answer [Code] [Data format] H2S highlights source-addressable evidence from a block-structured document,
Hugging Face Datasets2026 · Image
PhysAlignPhysAlign PhysAlign is a bilingual multimodal physics benchmark for testing whether a model can make local observations and bind textual or visual references to the correct candidate entity. It does not ask the evaluated model to solve the original examination problem. PhysAlign was constructed from selected examples in previously released public benchmarks, followed by source adaptation, probe ge
Hugging Face Datasets2026 · Table · Parquet
US Constitution Annotated (Congressional Research Service)US Constitution Annotated (Congressional Research Service) The Constitution of the United States of America: Analysis and Interpretation, which the Congressional Research Service (CRS) prepares and the United States Congress prints as a Senate Document: every edition and supplement that GovInfo holds, with the full text of each part of the book and GovInfo's record of it. Nothing here is edited by
Hugging Face Datasets2026 · dataset
TUSI-BenchTUSI-Bench TUSI-Bench is a Persian legal benchmark designed for research on legal question answering, legal knowledge retrieval, and evaluation of large language models in the domain of Iranian law. TUSI-Bench consists of three complementary components: TUSI-KB: a structured Persian legal knowledge base containing legal provisions from Iranian law. TUSI-QA: an open-ended Persian legal question-ans
Hugging Face Datasets2026 · Text
SPIRAL-Bench v0SPIRAL-Bench v0 A small benchmark for testing whether the wording of a retrieval query changes the balance of the evidence a retriever returns. Built for the Ouroboros project, which studies self-confirming retrieval loops in agentic RAG. The question this dataset exists to answer In agentic RAG, the system writes its own follow-up search queries, and it writes them using what it currently believe
Hugging Face Datasets2026 · Text
abute-21/amharic-geez-numerical-blindspotBlind Spots of Frontier Models: Tokenizer-Induced Numerical & Temporal Collapse in Ge'ez and Amharic Author: Teshome Birhanu Cheru Affiliation: Addis Ababa University, Electrical and Computer Engineering Target Fellowship: Fatima Fellowship 2026 Technical Challenge Evaluated Model: Qwen/Qwen2.5-3B-Instruct (3B Parameters) Artifacts Repository: Evaluation Benchmark (benchmark_ethiopian_reasoning.js
Hugging Face Datasets2026 · Table · Parquet
FineEnvs/SmolDataEnvs📈 SmolDataEnvs 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest. Data-analysis tasks as a plain, load-and-go dataset: no runtime, no framework required. Each row is one self-contained ta
Hugging Face Datasets2026 · Table · Parquet
BioDecision SFT v2.2BioDecision SFT v2.2 1.1M biomedical decisions in one format: a source text, a question, lettered options, one correct letter. Built to train BioDecision-4B with Together AI's Tev1 recipe; usable with any classifier or LLM that scores options. Split Rows Tokens Use train 1,080,373 420,677,060 training dev 12,235 4,768,045 model selection calibration 12,267 4,806,484 temperature fitting only benchm
Hugging Face Datasets2026 · dataset
MacJev-0.8B Decision Training Data (round 2)MacJev-0.8B decision training data (round 2, 2026-09-24) This repository holds the data used to train MacJev-0.8B. MacJev-0.8B is a Qwen3.5-0.8B decision model. For each option it outputs a score, read as h[slot] · (w_yes − w_no) at a -> slot placed after that option. It was trained on inputs of up to 25,600 tokens. The repository also holds the evaluation results of the one-H100 training run. Lay
Hugging Face Datasets2026 · Text
Fable 5.1'ed SFT DataDataset Description This dataset contains 473,635 agentic coding and reasoning high-quality multi-turn traces originating from the Step 3.5 Flash SFT Code dataset. It was remade to sound and act very similar to the Fable 5.1 model on max reasoning effort in Fable-5.1-Max-Reasoning-Filtered-10000x. It holds over 2,000,000,000 tokens of step-by-step chain-of-thought programming across multiple compl