Hugging Face Datasets2026 · Text
SecondState FAB — Finance Agents BenchmarkFAB — Finance Agents Benchmark FAB is an open-source project for benchmarking LLM agents' ability to perform financial due diligence in a synthetic company data room. FAB consists of a dataset of tasks containing agent instructions, documents and rubrics, and an execution harness for running and evaluating agents. This repository contains the dataset; the harness is available on GitHub. Dataset 50
Hugging Face Datasets2026 · dataset
Paderborn LDM run-to-failure ball bearings (time-varying load and speed), int16Paderborn LDM run-to-failure bearings, lossless int16 repack Repack of Run-to-failure data set of ball bearings subjected to time-varying load and speed conditions (O. K. Aimiyekagbon, Paderborn University, Zenodo, DOI 10.5281/zenodo.10805042, CC BY 4.0). Nothing is resampled or filtered: each accelerometer value is the integer code of the data logger's grid, and volts = offset[c] + step[c] * code
Hugging Face Datasets2026 · Image
XiaomiMiMo/MiMo-V2.6-RL-ossAgentic RL Environments RL training environments for LLM agents. Domain Task Family Verifier Code Software engineering Executable tests Cyber Vulnerability reproduction Rule checks General Knowledge work Rubric-based judging Visual Web development Visual grading Music Symbolic music composition Rule checks Docker images: https://hub.docker.com/r/xiaomimimo/mimo-v2.6-rl-oss Training code: https://g
Hugging Face Datasets2026 · Text
Bourse Assistant RL DatasetBourse Assistant RL Dataset این مخزن دادهها را برای پروژه دستیار بورس [لینک پروژه] منتشر میکند. مجموعهدادهای برای فاینتیون با روش یادگیری تقویتی یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعهای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند. محتوای مجموعهداده (Dataset Summary) داده به دو
Hugging Face Datasets2026 · Text
Bourse Assistant SFT DatasetBourse Assistant SFT Dataset این مخزن دادهها را برای پروژه دستیار بورس [لینک پروژه] منتشر میکند. مجموعهدادهای برای فاینتیون نظارتشده یک دستیار هوش مصنوعی در حوزه بورس اوراق بهادار ایران. هر نمونه شامل مجموعهای از خبرهای مربوط به یک نماد بورسی مشخص در یک تاریخ معین است، و مدل باید روند قیمت (مثبت/منفی) را همراه با توضیح تخمین بزند. محتوای مجموعهداده (Dataset Summary) داده به دو زیرمجموعه ت
Hugging Face Datasets2026 · Table · Parquet
FinancialAuditBenchFinancialAuditBench Code · Trajectories FinancialAuditBench evaluates AI agents on financial statement audit tasks. It contains 90 tasks across 6 synthetic audit engagements in manufacturing and staffing services. Agents use supporting documents to complete or review spreadsheet workpapers. Task settings Configuration Tasks Starting workpaper Goal completion 45 Blank template Perform the specified
Hugging Face Datasets2026 · Text
biopharma-benchBiopharma Bench V0.1 Biopharma Bench is an evaluation suite measuring whether frontier AI agents can perform professional, regulated desk work in biopharmaceutical drug and medical device development. Instead of testing models on isolated multiple-choice questions or pre-packaged single-document summaries, Biopharma Bench places agents into realistic employee seats across complete corporate operat
Hugging Face Datasets2026 · Text
OEP-BenchOEP-Bench OEP-Bench (Orex Enterprise Policy Benchmark) is a synthetic benchmark for evaluating retrieval-augmented generation (RAG) and enterprise knowledge systems under policy versioning, access-control, citation, temporal-reasoning, and security constraints. OEP-Bench is built around NovaCore Technologies, a fictional enterprise used solely as the organizational setting for the benchmark. NovaC
Hugging Face Datasets2026 · dataset
Infoton/Infoton_Virtual_Mitochondria_Picard_Validation_ResultsHugging Face Datasets2026 · Text
WildSongBench🤗 WildSongBench A benchmark for full-song music generation 192 prompts · 94 Chinese · 98 English 🎵 YuE2 project · 🚀 Quick start · 📊 Benchmarks · 🔁 Reproduce · 📄 arXiv · PDF · 📚 Citation WildSongBench (WSB) contains 192 song-generation prompts: 94 Chinese and 98 English, used in the YuE2 benchmarks. This repository pr
Hugging Face Datasets2026 · dataset
Infoton/Infoton_Virtual_Cell_Long_Covid_SARS-COV-2Hugging Face Datasets2026 · dataset
JDC/UNHCR FCV Data-Use Corpus (91 documents)JDC/UNHCR FCV Data-Use Corpus (91 documents) Annotated document corpus behind the FCV data-extraction paper (jdc_unhcr_fcv_data_extraction_paper.tex): 91 documents spanning four collections — UNHCR/ReliefWeb field reports (43), SEIS & humanitarian briefs (25), World Bank Policy Research Working Papers (14), and World Bank Project Appraisal Documents / PADs (9). Structure data/corpus.jsonl — one ro
Hugging Face Datasets2026 · Text
FinFIRSTFinFIRST: Financial Information Retrieval, Sourcing and Traceability Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China
Hugging Face Datasets2026 · Image
Synthetic Medical Document Recognition BenchmarkSynthetic Medical Document Recognition Benchmark This dataset contains synthetic, English-language medical records rendered as documents for evaluating automated data extraction and de-identification systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple visual representations derived from that record. Every rendered document is clearly marked as synthetic. This makes the d
Hugging Face Datasets2026 · dataset
markov-ai/cad-1000-hoursCAD-1K Open v2 - 1,018.1229 Hours 509 end-to-end, single-display Windows CAD task recordings across seven CAD software families. Each task contains: task_desc.json - task prompt, application, reference-input paths, and expected deliverables input_files/ - reference inputs named input.ext or input_N.ext output_files/ - submitted CAD deliverables and supplemental outputs named output.ext or output_N
Hugging Face Datasets2026 · Table · Parquet
datalab-to/omni_extract_benchOmni Extract Bench We weren’t satisfied with the current benchmarking options for extraction. They were biased, didn’t use realistic data and were hard to audit. Our view is that an extraction benchmark should do two things: Help customers choose the right vendor; and Give engineers a way to diagnose what’s actually going wrong in a given model. That’s why we built OmniExtractBench. OmniExtractBen
Hugging Face Datasets2026 · Image
ExtractBenchExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated rec
Hugging Face Datasets2026 · Image
lenamerkli/distilled-webDataset Card for lenamerkli/distilled-web This dataset consists of web-scraped data using a custom crawler purpose-built for each website. Dataset Details Dataset Sources Repository: https://github.com/lenamerkli/distilled-web Uses This dataset is useful for training large language models. The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction
Hugging Face Datasets2026 · Image
brmiller/brain-microprotein-atlasHugging Face Datasets2026 · Table · CSV
GeneBench-ProGeneBench-Pro Public Case Studies This repository contains public GeneBench-Pro case studies. It is the self-contained package intended for public distribution, including Hugging Face publication. Package Layout <repo-root>/ ├── .gitattributes ├── README.md ├── LICENSE ├── problems.csv ├── checksums.sha256 ├── manifest.json ├── reference_definitions.md ├── reference_grader.py └── problems/ └── <ev
Hugging Face Datasets2026 · Image
ECCV 2026 CAD Challenge DataECCV 2026 CAD Challenge Data This challenge is part of the workshop The Path to Manufacturing: Evolving 3D Generation to Intelligent Computer-Aided Design. Workshop homepage: https://3dgen-cad-workshop.github.io/ Challenge submission Space: https://huggingface.co/spaces/jingwei-xu-00/eccv2026-cad-challenge Dataset rendering and preparation code (only .step files are required): https://github.com/D
Hugging Face Datasets2026 · Table · CSV
GeneBench-ProGeneBench-Pro Public Case Studies This repository contains public GeneBench-Pro case studies. It is the self-contained package intended for public distribution, including Hugging Face publication. Package Layout <repo-root>/ ├── .gitattributes ├── README.md ├── LICENSE ├── problems.csv ├── checksums.sha256 ├── manifest.json ├── reference_definitions.md ├── reference_grader.py └── problems/ └── <ev
Hugging Face Datasets2026 · Text
PrivacyBenchPrivacyBench PrivacyBench is a benchmark for de-identifying semi-structured data exports from work tools like email, messaging, shared file storage, calendar, and so forth. The benchmark focuses on the identification and synthesis of PII in the unstructured text fields of the data export, and introduces novel metrics for evaluating synthesis quality. This version (v2) contains one month of workspa
Hugging Face Datasets2026 · Table · Parquet
RedlineBenchAbstract Crosby–micro1 RedlineBench measures contract negotiation as a sequence of judgment calls rather than a collection of isolated clause edits. It captures multi-turn redlining workflows through simulations grounded in realistic SaaS transactions and attorney-generated explanations of key redline decisions, and evaluates models across five dimensions: legal correctness, commercial alignment,
Hugging Face Datasets2026 · Image
PureDocBenchMain Leaderboard 58 models · 3 matched tracks · 🏆 Search, filter & sort the leaderboard → The current evaluation covers 13 pipeline / multi-stage specialists, 19 end-to-end specialists, and 26 general-purpose VLMs. Each track contains 1,475 pages. The top 10 by the three-track mean, Avg₃, are shown below. Rank Model (release) Type Clean ↑ Digital ↑ Real ↑ Avg₃ ↑ 1 GLM-5.3-Flash (2026-08) General V