{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:reinforcement-learning",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Text · gated
Dissei Financial Judgment — Full Evaluation PackageDissei Financial Judgment — Full Evaluation Package Harbor · GitHub · Website · Hugging Face Dissei explores financial reasoning and analysis behind institutional investment and credit decisions: interpreting evidence, weighing trade-offs and reaching well-supported conclusions. Financial Judgment contains seven tasks drawn from one completed private-equity deal. Each asks for a written analytical
Hugging Face Datasets2026 · Text
JARVIS Autonomous Reasoning & DPO Dataset🧠 JARVIS Autonomous Reasoning & DPO Trajectories Curated & Engineered by Boss Ratan (Al-Amin Ahmed Ratan) Welcome to the official repository of verified synthetic reasoning traces and Direct Preference Optimization (DPO) pairs, generated autonomously by the JARVIS Mark-LIV Autonomous System. 📊 Dataset Structure Each sample contains verified multi-step reasoning traces formatted for state-of-the-ar
Hugging Face Datasets2026 · Table · Parquet
MiroShark Social + Prediction Market SimulationMiroShark Social + Prediction Market Simulation Agent decisions from MiroShark simulations (GitHub). In each simulation, LLM agents with distinct personas (companies, founders, communities, regulators, commentators) share a Twitter/Reddit-style feed and a Polymarket-style prediction market. Every round, each agent reads the feed (or its portfolio and the open markets) and decides what to do: post,
Hugging Face Datasets2026 · Table · Parquet · gated
Datapoint Text-to-Video Human Preferences (326K)Text-to-video human preferences: 326K votes across 15 models This dataset contains the complete voting record behind the Datapoint Video Bench leaderboard: 325,520 validated pairwise votes — exactly 10 for each of 32,552 video pairs. The votes compare 15 text-to-video models on 314 prompts built to stress motion, physics, and temporal consistency, judged by 22,982 annotators in 187 countries. Ever
Hugging Face Datasets2026 · Table · Parquet
Jev Decisions v1Jev Decisions v1 12M canonical agent-decision records for tool selection, routing, value prediction, completion, and local agent control. Jev Decisions v1 is a derived, decision-oriented corpus built from public agent trajectory datasets. It canonicalizes heterogeneous trajectories into a shared learning interface: state + available candidate decisions -> target / outcome / eligibility mini-Jev is
Hugging Face Datasets2026 · Table · CSV
Agent Memory Resilience & Poisoning BenchmarkAgent Memory Resilience & Poisoning Benchmark Dataset Summary This benchmark dataset evaluates resilience, negative transfer, and memory poisoning mitigation in autonomous LLM agent architectures (such as LangGraph, AutoGen, and CrewAI). When autonomous agents record distilled self-reflections after attempting tasks, external stochastic failures or subtle API deprecations often cause agents to com
Hugging Face Datasets2026 · Table · CSV · gated
ARC-AGI-3 Schema Gameplay Trajectories (GPT-5.6 Sol)ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol This release contains every gpt-5.6-sol gameplay trajectory produced on our cluster with the world_model_v5 agent harness — 100 runs across the 25 public ARC-AGI-3 games — plus a dependency-free scoring utility. It is the GPT-5.6 Sol member of a family built by the same harness and the same sanitizer, so trajectories can be compared game by game
Hugging Face Datasets2026 · Table · CSV · gated
ARC-AGI-3 Schema Gameplay Trajectories (Claude Opus 4.8)ARC-AGI-3 Schema Gameplay Trajectories — Claude Opus 4.8 This release contains the best claude-opus-4-8 / max trajectory for each of the 25 public ARC-AGI-3 games, plus a dependency-free scoring utility. It is the Opus 4.8 counterpart of arc-agi-3-schema-traces-fable5, produced by the same agent harness (world_model_v5) and the same sanitizer, so the two collections can be compared game by game. E
Hugging Face Datasets2026 · Table · Parquet
PrimeIntellect/Recursive-Task-SynthesisRecursive Task Synthesis Tasks without completed platform artifacts or with unresolved VM validation failures are temporarily excluded. exclusions.json records the exact IDs, reasons, build IDs where available, and evidence dates/runs. Exclusions affect both metadata rows and complete TAR task packages. Runtime failures are not image-build failures or proof of incorrect gold solutions. This filter
Hugging Face Datasets2026 · dataset
Yunncheng/gamewam-minecraftGameWAM Minecraft Datasets This repository contains three Minecraft gameplay datasets used by GameWAM: A World Action Model for Video Games. All three use LeRobot v2.1 format and share a 22-D keyboard/mouse action space, a 6-D raw state (pitch, yaw, cursor x/y, hotbar, and GUI-open state), and a 15-D canonical proprioceptive representation. Project page: https://yunncheng.github.io/GameWAM/Code: h
Hugging Face Datasets2026 · Table · CSV
sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024)
Hugging Face Datasets2026 · Table · Parquet · gated
Datapoint Text-to-Speech Human Preferences (315K)Text-to-speech human preferences: 315K votes across 15 models This gated dataset contains the evaluation record behind Datapoint Audio Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech models in a complete round-robin over 300 English prompts. The prompt set covers eight practical voice-agent categories, and every generated sample is included as a typed audio record. The source ev
Hugging Face Datasets2026 · Table · Parquet
LITCOIN Proof-of-Research CorpusLITCOIN Proof-of-Research Corpus 191,484,662 AI research submissions, produced by 81,224 anonymous contributors and 470 model variants competing against each other, every row executed in a sandbox and scored. This is the complete output of the LITCOIN protocol, which ran on Base from March to August 2026. Autonomous AI agents were paid in a permissionless token to solve real optimization problems
Hugging Face Datasets2026 · dataset
EmbodiedCity/ANWM-DatasetANWM-Dataset Training / evaluation trajectories for ANWM (Aerial Navigation World Model), released with the paper Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space. Code: https://github.com/EmbodiedCity/ANWM.code Model: EmbodiedCity/ANWM Contents Sharded tar archives of AirVLN-16 style trajectories (airvln_16-*.tar). Each archive preserves the original folder layout
Hugging Face Datasets2026 · Text
osunlp/early-experienceEarly Experience — Reproduction Data Supervised fine-tuning data for reproducing Agent Learning via Early Experience across 8 agent environments. Each environment provides data for three training paradigms: IL — Imitation Learning: expert SR — Self-Reflection: expert + reflection IWM — Implicit World Modeling: iwm (world model) → expert Code: OSU-NLP-Group/EarlyExperience Usage from datasets impor
Hugging Face Datasets2026 · Image
SVG Generation Benchmark (Static)Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,355,161 human responses, collected with the Rapidata Python SDK, comparing how well 30 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
Hugging Face Datasets2026 · dataset · gated
lightcone02/OmniContact-DatasetOmniContact Dataset: Contact-Rich Humanoid Object Interaction Project Page OmniContact contains human-object interaction motion capture, processed Unitree G1 trajectories, and motion clips for contact-rich box manipulation and soccer interactions. This release contains 690 processed source trajectories and 2,231 clips. Every source NPZ has exactly one retained raw mocap capture and one portable BV
Hugging Face Datasets2026 · Image
SVG Generation Benchmark (Static)Rapidata Static SVG Generation Benchmark Built by Rapidata. This dataset contains 1,918,367 human responses, collected with the Rapidata Python SDK, comparing how well 42 frontier LLMs generate static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
Hugging Face Datasets2026 · Table · Parquet
Codex 5.5 Handwritten TropicalGT/ToricGT ToT ReasoningCodex 5.5 Handwritten TropicalGT/ToricGT ToT Reasoning This dataset contains individually authored Tree-of-Thought style reasoning records for TropicalGT/ToricGT training. Each row represents one accepted JSON reasoning artifact, flattened into stable Parquet columns for filtering and dataset-viewer compatibility while preserving the complete canonical JSON object in record_json. The source datase
Hugging Face Datasets2026 · Table · Parquet
ARC-AGI-3 World Model TracesARC-AGI-3 World Model Traces This dataset contains ARC-AGI-3 transition traces in the same parquet schema used by HHazard/arc-agi-3. Each row is one environment transition: state, game_id, level_id, action_id, action_args, next_state, level_done, frame_idx, origin, transformation, player state and next_state are 64x64 ARC grids stored as nested integer arrays. action_id is the ARC-AGI-3 action kin
Hugging Face Datasets2026 · Text
nvidia/Nemotron-RL-Ultra-Training-BlendsDataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Ea
Hugging Face Datasets2026 · dataset
SETA-EnvSETA-Env SETA-Env is an open-source verifiable RL terminal environment dataset for community training and evaluation. This release contains two top-level subsets: SETA_Synth: synthesized tasks SETA_Evolve: evolved variants of terminal-agent tasks The current release contains 4567 environments: SETA_Synth: 3255 SETA_Evolve: 1312 What Is Included Each task is packaged as a self-contained Harbor-styl
Hugging Face Datasets2026 · dataset
DecodingTrust-Agent Platform (DTAP-BENCH)DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the full collection of the agent trajectories produced from evaluating the DTap-Bench from DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Da
Hugging Face Datasets2026 · dataset
DecodingTrust-Agent Platform (DTAP-BENCH)DecodingTrust-Agent Platform A Controllable and Interactive Red-Teaming Platform for AI Agents. This is the per-task dataset for the DecodingTrust-Agent Platform (DTAP), spanning 14 real-world domains and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task ships the configuration the evaluator need
Hugging Face Datasets2026 · dataset
exploitbench/v8ExploitBench V8 — v8-codex-ace-83a40e1-ptf81548b Per-cell exploitation results from the V8 JavaScript engine benchmark, with full transcripts, tool-call logs, and capability grading. This dataset is the academic record for ExploitBench: succeeded runs and model-failed runs both ship, including cells where the model gamed the grader (see audit.json). Envs in this revision 41 environments. Full list