Hugging Face Datasets2026 · dataset
Areyliu/ManiEdit-statsThis repository contains the precomputed second-moment statistics for the paper [ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective]. These statistics are used to evaluate sequential unstructured knowledge editing in large language models. For code and usage instructions, see the GitHub repository. License: Apache 2.0.
Hugging Face Datasets2026 · Table · Parquet
rishavk77/amc26-entity-resolution-embeddings-testAMC26 — TEST-Split Business Entity Resolution Embeddings (multilingual-e5-base) Precomputed sentence embeddings for the test split of the Amazon ML Challenge 2026 business entity-resolution task. Same pipeline, same model and same channels as the train-split release, so the two are directly comparable. Purpose: blocking / candidate retrieval. For each Source 1 business, narrow the ~10M-record pool
Hugging Face Datasets2026 · dataset
FineEnvs/MiMo-V2.6-RL-harbor-cyberMiMo-V2.6-RL Cyber (Harbor) Reproduce a real memory-safety crash (ARVO). 1,000 Harbor tasks from the Cyber domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. The agent gets a sanitizer report and the project's source and fuzzer binary, and has to submit an input that crashes it in the expected function with the
Hugging Face Datasets2026 · dataset
HaomingLuo/AgentFEM-Controlled-Transient-HeatAgentFEM T3 Controlled Transient Heat This dataset contains 192 complete finite-element trajectories from 64 independent physical scenarios. A two-dimensional metallic solid is heated by three localized volumetric sources and cooled by convection along its exterior. Every scenario branches from the same state at 300 s into three possible futures: maintain source power, throttle all sources to 40%,
Hugging Face Datasets2026 · dataset
FineEnvs/MiMo-V2.6-RL-harbor-codeMiMo-V2.6-RL Code (Harbor) Fix a real issue in a real repository. 2,698 Harbor tasks from the Code domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. The agent gets an issue and the repository at its base commit. After it finishes, the files the hidden test patch touches are reset, the patch is applied, and the
Hugging Face Datasets2026 · Table · Parquet
orcarouter/OrcaJev-DecisionsOrcaJev Typed Decisions — Full Corpus One problem, one System-One decision: state + typed question → typed answer + calibrated probability, in a single forward pass, no decode loop. Each case holds exactly one question, keyed q, of type choice, noul, or score. Gold is always verifiable or human/dataset-annotated — never model-generated. At a glance Cases 1,081,380 · train 1,069,380 / v
Hugging Face Datasets2026 · Text
ARC-AGI-2 masked-demonstration episodes (JEPA-ready)ARC-AGI-2 masked-demonstration episodes (JEPA-ready) A task-level reasoning dataset built from the official ARC-AGI-2 tasks. Every row is an episode: a set of demonstration pairs (context), one held-out test_input, and the target_output the model must produce by inferring the rule shared by the demonstrations. It is not an input grid -> output grid dataset; the unit of learning is the task rule. B
Hugging Face Datasets2026 · Table · Parquet
Maia3 Training DataMaia3 Training Data Preprocessed chess position data for training Maia3 human move-prediction models. Positions are stored as precomputed features, not raw FENs, so training reads and expands them without re-tokenizing every epoch. Contents path rows size lichess_parquet/train_YYYY-MM_precomputed.parquet (31 files, 2023-01 … 2025-07) 334,438,119 ~33 GB allie_data/test_precomputed.parquet 884,049 6
Hugging Face Datasets2026 · dataset
FineEnvs/SmolDataEnvs-harbor-train📊 SmolDataEnvs: Harbor (train) 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest. The training suite: 5,000 hands-on data-analysis tasks. Each one drops an agent into a sandbox with a rea
Hugging Face Datasets2026 · dataset
AgentFEM Layered Thermoelastic Bending 2DAgentFEM Layered Thermoelastic Bending 2D This dataset contains 512 verified two-dimensional finite-element simulations of perfectly bonded bilayer and trilayer strips under a uniform positive temperature increment. It supports scalar surrogate modeling, mesh-based field learning, scientific-ML benchmarking and reproducible CAE workflow studies. Dataset summary 256 bilayers and 256 trilayers. 256
Hugging Face Datasets2026 · dataset
AgentFEM Material Loading MemoryAgentFEM Material Loading Memory An open, reproducible research dataset for path-dependent material modeling, neural constitutive surrogates and finite-element deployment tests. The repository contains six staged, explicitly separated releases. Start here Goal Recommended entry Understand the current dataset This page and the sealed-test data card Train or benchmark a constitutive model DENIM star
Hugging Face Datasets2026 · Image
China Universities DatasetChina Universities Dataset 🎓 582 Chinese universities — names (English + 中文), city, province, category, 985/211/双一流 tags, logos, and ShanghaiRanking (软科中国大学排名) 2026 national ranks, scores, and institution-page links. Shipped as ready-to-use JSON · CSV · SQL, no scraping needed. One import and you have a structured Chinese university database: searchable slugs, province→city geography, prestige tag
Hugging Face Datasets2026 · Text
biopharma-benchBiopharma Bench V0.1 Biopharma Bench is an evaluation suite measuring whether frontier AI agents can perform professional, regulated desk work in biopharmaceutical drug and medical device development. Instead of testing models on isolated multiple-choice questions or pre-packaged single-document summaries, Biopharma Bench places agents into realistic employee seats across complete corporate operat
Hugging Face Datasets2026 · Image
Revealing Item-level Differences of Generalist agents across digital and physical EnvironmentsRIDGE: Revealing Item-level Differences of Generalist agents across digital and physical Environments Item-level results and agent traces from Astra, Opus 5.5, and other Frontier Models Demonstrate Jagged Performance Across SoTA Agentic Tasks from Web Browsing to Robotics. We ran seven frontier models on web browsing and three on driving, tabletop manipulation, industrial procedures and toy assemb
Hugging Face Datasets2026 · dataset · gated
OSWorld-Science task dataOSWorld-Science — task data Everything the OSWorld-Science benchmark (https://github.com/DiscoAILab/OSWorld-Science) needs besides the code, one directory per domain: the task definitions, the bytes the solver's virtual machine receives, the ground truth the host-side grader scores against, and the prepared VM image. <domain>/tasks/<task_id>.json task definitions (instruction, staging steps, deliv
Hugging Face Datasets2026 · Table · CSV
Gaze as Evidence for Common Grounding: MapTask and MUNDEX FeaturesGaze as Evidence for Common Grounding Processed, window-level gaze features for Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX by Nan Li, Albert Gatt and Massimo Poesio (MINT 2026). Paper on arXiv · Hugging Face paper page · GitHub: data and analysis code The dataset connects gaze measurements with reference-alignment annotations in MapTask and retrospective u
Hugging Face Datasets2026 · dataset
Synthetic Bedroom Layouts (ABO furniture)Synthetic Bedroom Layouts 4,000 procedurally composed bedroom layouts with 27,025 furniture placements, in the box format used by indoor scene-synthesis models (ATISS-style boxes.npz). Furniture is drawn from Amazon Berkeley Objects (CC BY 4.0), so the whole dataset is redistributable and usable commercially. Nothing here derives from 3D-FRONT, 3D-FUTURE, or any dataset that restricts redistributi
Hugging Face Datasets2026 · Text
INSIGHT-Bench v1INSIGHT-Bench v1 A human-curated object-goal navigation benchmark: 1,097 episodes over 210 scenes, each episode a short natural-language instruction, a start pose, a goal position and a success radius, defined on Z-up, metre-scaled USD conversions of four scene sources -- HM3D, Matterport3D, InteriorGS and Habitat-GS (3D Gaussian Splatting). It is evaluated in NVIDIA Isaac Sim by the INSIGHT-Bench
Hugging Face Datasets2026 · Table · Parquet
Infoton Identity — 735 Human ProteinsAn Infoton Dose of RxRx — 735 Human Proteins Version: 1.0.0 Released: 2026-09-06 Author: Walker, January Natatia — Infoton DOI: 10.5281/zenodo.18210355 ORCID: 0009-0000-6843-2051 Overview The Infoton Identity Physics Engine computes in seconds per protein and aligns with what cellular imaging measures. The coordinates in the An Infoton Dose of RxRx dataset were derived before correlation was run w
Hugging Face Datasets2026 · Table · Parquet
DeepSeek-V4-Pro 0813 AgenticDeepSeek-V4-Pro 0813 Agentic (DS4) A standalone, verifiable-first agentic training corpus: 19,072 training traces plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families, each row admitted only after passing a deterministic programmatic verifier. The corpus is desig
Hugging Face Datasets2026 · Table · Parquet
Football Analytics SQLite Database⚽ Football Analytics Database Overview This dataset contains a processed SQLite database generated from StatsBomb Open Data. The objective of this dataset is to provide a ready-to-use relational database for football analytics, eliminating the need to parse and transform thousands of raw JSON files. The database was created as part of the Football Analytics Dashboard project and is intended for: F
Hugging Face Datasets2026 · Table · Parquet
GitSkillsGitSkills: A Dataset of Agent Skills on GitHub Paper (arXiv:2608.10906) · Sample repository · Zenodo DOI: 10.5281/zenodo.21875637 An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent loads the skill when it judges that a task matches the skill description. Anthropic introduced the format
Hugging Face Datasets2026 · dataset · gated
SciAgentArena — Terminal-Bench Benchmark (self-contained)SciAgentArena — Terminal-Bench (Harbor) Benchmark 304 agent tasks across seven scientific benchmarks, with all data included. Clone it, point tb run at a task, and it works — no other download required. Tasks use the shipped benchmark evaluators, including the drug-discovery repairs and EHR T2 value checks documented in tasks/GENERATION.md. Benchmark Tasks Data Scorer Drug discovery (C1–C5) 75 cur
Hugging Face Datasets2026 · Table · Parquet
THBKG — Temporal Heterogeneous Biomedical Knowledge GraphTHBKG — Temporal Heterogeneous Biomedical Knowledge Graph A dated biomedical knowledge graph built from Open Targets 26.03 (with Reactome, ChEMBL and ClinicalTrials.gov), plus a clinical-advancement benchmark: rank target–disease pairs by their likelihood of advancing to Phase III, scored only from evidence datable strictly before each pair's decision year. Every temporal edge carries the year its
Hugging Face Datasets2026 · dataset
The Open Distillation Codex📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tut