Hugging Face Datasets2026 · dataset
WideSWEWideSWE: Can Coding Agents Coordinate Changes Across Repositories? A benchmark for implementing one shared feature or bug fix across multiple repositories. Overview WideSWE contains 120 real-world software-engineering tasks across 41 software ecosystems, balanced between 60 bug fixes and 60 features. Each task requires coordinated changes across multiple repositories in the same ecosystem. Agents
Hugging Face Datasets2026 · Text
KaliBench-VerifiedDataset Card for KaliBench KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards (NeurIPS 2026 Evaluations and Datasets Track) Authors: Pengfei Li1*, Naufal Suryanto1*, Sicheng Zhang1, Muzammal Naseer1,2 1Khalifa University, 2University of Western Australia *Equal contribution 💻 GitHub Code | 📊 Dataset Dataset summa
Hugging Face Datasets2026 · Text
Nickyang/RazorCalRazorCal RazorCal is a general-purpose calibration corpus released with RAZOR. RazorCal.json contains 2,048 samples across seven domains (about 8 MB). Each record includes input messages and source attribution. The records carry no architecture-specific fields, so the corpus also suits FFN or layer pruning, perplexity measurement and other calibration-based compression; RAZOR uses it to estimate p
Hugging Face Datasets2026 · Text
GSM8K-UZ (Cyrillic)GSM8K-UZ (Cyrillic) Uzbek GSM8K in the Cyrillic script. Part of a parallel pair: kurbanovxurshidbek/gsm8k-uz-lat and kurbanovxurshidbek/gsm8k-uz-cyr. Both contain exactly the same problems (same idx), differing only in script. Split Rows train 7417 test 1308 Source and construction Latin text: NeuronUz/gsm8k-uz, a machine translation of openai/gsm8k (main) into Uzbek Latin. Cyrillic text: automati
Hugging Face Datasets2026 · Text
The First And Best Fable 5.1 Reasoning DataDataset Description This dataset contains 10,000 agentic coding and reasoning multi-turn high-quality traces generated by the new Fable 5.1 model using max reasoning effort. It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains. It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces
Hugging Face Datasets2026 · Text
GSM8K-UZ (Latin)GSM8K-UZ (Latin) Uzbek GSM8K in the Latin script. Part of a parallel pair: kurbanovxurshidbek/gsm8k-uz-lat and kurbanovxurshidbek/gsm8k-uz-cyr. Both contain exactly the same problems (same idx), differing only in script. Split Rows train 7417 test 1308 Source and construction Latin text: NeuronUz/gsm8k-uz, a machine translation of openai/gsm8k (main) into Uzbek Latin. This dataset reproduces the L
Hugging Face Datasets2026 · Text
Agentic Coding Chain-of-Thought Dataset🤖 Agentic Coding CoT Dataset A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with e
Hugging Face Datasets2026 · Table · Parquet
Qwen3.8-Max DistillationQwen3.8-Max Distillation A quality-filtered derivative of Qwen3.8-Max Distillation 50K by r0b0tlab, prepared for local training and fine-tuning on consumer hardware. This repository takes the original 49,772-example dataset and produces a substantially smaller training set focused on coding, reasoning, instruction following, and tool use. [!CAUTION] Terms and provenance notice — not cleared for un
Hugging Face Datasets2026 · Text
BaRe-Mem DataBaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation Overview In multi-agent systems, a central model can consult advisors, but advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. BaRe-Mem is an online Bayesian reliability memory for multi-agent consultation: it estimates each advisor's reliability fr
Hugging Face Datasets2026 · Table · Parquet
Whittle distillation set (Qwen3.8-27B reasoning traces + top-20 logprobs)Whittle distillation set: Qwen3.8-27B traces with top-20 logprobs About 25,000 complete, quality-filtered answers from Qwen3.8-27B (thinking on), each with the teacher's top-20 log-probabilities at every answer token. It is built for logit-level knowledge distillation: a student can match the teacher's full per-token distribution, not just its sampled text. We use it to train Whittle-Qwen-3.8-35B-
Hugging Face Datasets2026 · dataset
Ali Data — Urdu & Roman Urdu CorpusAli Data — Urdu & Roman Urdu Corpus A large public-domain/open collection of Urdu (اردو) and Roman Urdu text for training Urdu language models. Stats 21,846,033 unique rows (deduplicated by exact text) Format: JSON Lines — each line: {"text": "...", "source": "..."} Scripts: Urdu (Arabic script) + Roman Urdu (Latin script) Sources Source Rows (approx) Hugging Face open datasets (news, QA, sentimen
Hugging Face Datasets2026 · Text
TozAI/Toz-1-SFTTöz-1 SFT verisi Töz-1 sohbet modelini eğitmek için kullanılan Türkçe SFT (talimat/sohbet) verisi: 12,297 örnek, 14,316 soru-cevap turu. Verinin neredeyse tamamı kodla üretildi (uretim/ klasöründeki scriptler). Amaç modele bilgi değil biçim öğretmek: 195M parametreli bir modele SFT'de bilmediği bir bilgiyi öğretmek uydurmayı öğretir (Gekhman ve ark., 2024). Bu yüzden: Bilgi soruları sadece doğru c
Hugging Face Datasets2026 · Table · CSV
Discursos del Congreso de los DiputadosDiscursos del Congreso de los Diputados (muestra balanceada) Descripción Este dataset contiene una muestra de 1088 intervenciones en el pleno del Congreso de los Diputados de España, entre 1996 y 2023. Se ha obtenido a partir del corpus ParlLawSpeech mediante un proceso de limpieza y un muestreo balanceado por año y partido político. Se ha creado como parte de la práctica de la asignatura Descubri
Hugging Face Datasets2026 · Table · Parquet
Satori JSX Layouts (FR)Satori JSX Layouts (FR) Dataset de paires (description en français → JSX Satori) pour l'entraînement de modèles capables de générer des mises en page visuelles (bannières, landing pages, cartes) sous forme de code JSX destiné à être rendu par Satori en SVG puis en image. Contenu Chaque exemple contient : prompt — description en français de la mise en page voulue (palette de couleurs, structure, co
Hugging Face Datasets2026 · Text
AgentWeave Tool-Routing BenchmarkAgentWeave Tool-Routing Benchmark Benchmark cases and evaluation artifacts for AgentWeave, a deterministic pre-inference routing layer for tool-rich language-model systems. Contents data/benchmark_cases.jsonl — the canonical, normalized Dataset Viewer split. raw/evaluation/ — frozen heterogeneous evaluation artifacts retained as reproducibility evidence and intentionally excluded from the Viewer s
Hugging Face Datasets2026 · Table · CSV
Old Games Transcript90s Games Transcript A small English-language corpus of narrative text from classic PC games. This dataset was assembled as a compact research corpus for game studies, digital humanities, discourse analysis, narrative analysis, computational stylistics, and computationally assisted close reading. Dataset Config Records Content caesar3 20 Mission briefings + victory messages diablo2_lod 7 Cinematic
Hugging Face Datasets2026 · dataset
Delta DatasetThe Delta Dataset V2 The Delta Dataset is a collaborative, open-source dataset created to help build and improve the next generation of AI models. The project is developed by Pyra Labs and can be used to train and fine-tune Delta models, as well as other AI models and projects. Rather than being limited to a single model or architecture, the goal of The Delta Dataset is to provide an open space wh
Hugging Face Datasets2026 · dataset
Doctor-Patient Conversations — All Human Diseases (Opus 5.5)Opus-5.5 generated Doctor-Patient Conversations for All Human Diseases The sequel to nisten/opus-doctor-patient-conversations-all-human-diseases (Opus 4.8). Same disease list, same 20-key schema, same ChatML conversations — regenerated from scratch with Claude Opus 5.5, one agent per disease, and held to a much stricter bar. Covers every human disease listed on my previous work here: nisten/all-hu
Hugging Face Datasets2026 · Table · Parquet
UltraData-CodeUltraData-Code 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series | 📖 Tech Report (Coming Soon) | 🤗 UltraData-Code-L2 Classifier English | 中文 📚 Introduction UltraData-Code is a complete implementation of the UltraData L0-L4 tiered data management framework. It covers four code data states from L0 through L3, with each level corresponding to a distinct construction stage. The pipeline starts
Hugging Face Datasets2026 · Text
Chokepoint: Tool-Call Mediation Under Indirect Prompt InjectionChokepoint Benchmark Scenario suites and evaluation traces for Chokepoint, a validated harness for evaluating tool-call mediation defenses against indirect prompt injection (IPI) in LLM agents. Companion to the paper Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection (arXiv:2609.32691, cs.CR). Code: github.com/AnimeshShaw/cho
Hugging Face Datasets2026 · Table · Parquet
DranxX/corpus-cleaning-v1Text Cleaning Dataset (raw → clean) A paired dirty → clean text dataset for training text cleaning/denoising models. Format: 3 columns — lang (language code), raw (dirty text), clean (target clean text). Purpose Training a seq2seq model that transforms dirty text into clean text, covering: Character repair: corrupted characters (0x00, BOM, ^, ~), stray digits, broken spacing Boilerplate removal: j
Hugging Face Datasets2026 · dataset
dLLM PRM Gap Datasets and diagnostic artifactsdLLM PRM Gap · Datasets 💻 Code • 🤗 Collection • 📄 Paper This dataset repository contains the selected trajectory corpus and small diagnostic artifacts for dLLM PRM Gap, a controlled study of process reward model (PRM) guidance and outcome reward model (ORM) reranking in discrete diffusion reasoning. Public release. This dataset accompanies the arXiv paper and the dLLM P
Hugging Face Datasets2026 · Table · Parquet
Federal Reserve Beige BookFederal Reserve Beige Book Every edition of the Beige Book, the Federal Reserve's Summary of Commentary on Current Economic Conditions by Federal Reserve District, from the first, May 20, 1970, when the Board's pages call it the Redbook, to the latest. One row per section: the national summary, each of the twelve districts, and the one special report (May 18, 1983). In the words of the Board's Bei
Hugging Face Datasets2026 · Text
H2S-Research/Highlight-Then-SummarizeHighlight-Then-Summarize Highlight-Then-Summarize (H2S) is a compress-then-reason approach for long-context understanding. It makes evidence localization and information integration explicit before final-answer generation: long document + question -> evidence -> question-conditioned summary -> answer [Code] [Data format] H2S highlights source-addressable evidence from a block-structured document,
Hugging Face Datasets2026 · Table · Parquet
FaithformBenchFaithformBench Benchmark A of FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation (Cornish, Ghinassi, et al., 2026). This is the main benchmark of the paper: all results in the main text are computed on it. FaithformBench-Regex is the same benchmark with rule-based (regex) perturbations instead. Code: https://github.com/Ighina/FaithformBench Related dataset