{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:feature-extraction",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Text
PublicHearingLDSExtendedPublicHearingLDSExtended Dataset PublicHearingLDSExtended é um dataset estruturado para verificação de alegações (claim verification) e recuperação de evidências em audiências públicas da Câmara dos Deputados do Brasil. Ele estende o dataset original PublicHearingBR_LDS, preservando 100% dos seus documentos e metadados, adicionando uma camada estruturada de alegações (claims) por orador com evidên
Hugging Face Datasets2026 · Text
HUMMBL 40k Multi-Agent Wicked Problems & Coordination CorpusHUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet. Dataset Overview The dataset provides structured visibility into how autonomous agents navigate comple
Hugging Face Datasets2026 · Text
SwarmTraces publisher artifacts and Sev observationsSwarmTraces publisher artifacts Sev-4B v0.3.0 research update The v0.3.0 response-policy checkpoint improves authored policy decisions from 127/175 to 147/175 across 35 held-out source programs. At its calibration-selected alert threshold, it flags 14/140 permitted decisions, a 10% false-alert rate on this panel. 26/28 registered checks pass. The DNS diagnostic remains 21/32, with one repaired ans
Hugging Face Datasets2026 · Table · Parquet
rishavk77/amc26-entity-resolution-embeddings-testAMC26 — TEST-Split Business Entity Resolution Embeddings (multilingual-e5-base) Precomputed sentence embeddings for the test split of the Amazon ML Challenge 2026 business entity-resolution task. Same pipeline, same model and same channels as the train-split release, so the two are directly comparable. Purpose: blocking / candidate retrieval. For each Source 1 business, narrow the ~10M-record pool
Hugging Face Datasets2026 · Text
Sev security evidence collectionSev security evidence collection Sev-4B v0.3.0 research update The v0.3.0 response-policy checkpoint improves authored policy decisions from 127/175 to 147/175 across 35 held-out source programs. At its calibration-selected alert threshold, it flags 14/140 permitted decisions, a 10% false-alert rate on this panel. 26/28 registered checks pass. The DNS diagnostic remains 21/32, with one repaired an
Hugging Face Datasets2026 · Image
The Food API: Packaged Food & Beverage Products Dataset (Sample)The Food API: Packaged Food & Beverage Products Dataset Normalized packaged-food and beverage product data, keyed by barcode, for food-tech, nutrition apps and e-commerce. This repository is the free evaluation sample of The Food API: one record per product with the barcode, the full ingredient statement and its parsed tree, declared and precautionary allergens, the nutrition panel, package claims
Hugging Face Datasets2026 · Table · Parquet · gated
akshatbakshi/amazon-ml-challenge-2026Amazon ML Challenge 2026: Business Entity Resolution Dataset This repository hosts the official dataset for the Amazon ML Challenge 2026 — Business Entity Resolution Challenge, packaged for high-speed multi-server distributed training, EDA, and evaluation. Both the original TSV files and optimized Parquet files (zstd-compressed) are provided. 📊 Dataset Overview & Statistics Total records across tr
Hugging Face Datasets2026 · Image
The Food API: Packaged Food & Beverage Products Dataset (Sample)The Food API: Packaged Food & Beverage Products Dataset Normalized packaged-food and beverage product data, keyed by barcode, for food-tech, nutrition apps and e-commerce. This repository is the free evaluation sample of The Food API: one record per product with the barcode, the full ingredient statement and its parsed tree, declared and precautionary allergens, the nutrition panel, package claims
Hugging Face Datasets2026 · Table · Parquet
DeepSWE PRM training embeddings (Qwen3-8B, 8k)DeepSWE PRM training embeddings (Qwen3-8B, 8k) Frozen Qwen3-8B last-token-pooled embeddings (4096-d, float16) of every step of the DeepSWE training-pool rollouts: 405,919 steps · 4,701 trajectories · 113 tasks. This is the data the released DeepSWE PRM heads (tarsur385/deepswe-prm-heads-8k) were fine-tuned on. Embedded with preprocessing/deepswe/embed_shard.py at max_model_len 8192: the state is t
Hugging Face Datasets2026 · Table · Parquet
pvv5385/cve-embeddingsCVE Flat Table A flattened, tabular export of the official CVE List v5 (CVEProject/cvelistV5), one row per CVE record, built for embedding / semantic-search / classification experiments. Source data is public-domain CVE Program data (CVE Record Format 5.x). This export was generated on 2026-09-17 from the latest daily baseline release. Fields cve_id — CVE identifier state — record state (e.g. PUBL
Hugging Face Datasets2026 · Table · Parquet
Yigit-Karaman/open-jobs-dailyOpen Jobs Daily 🌍💼 Commercial vendors often charge upwards of $1,000/month for firehose access to global job market data. This dataset democratizes that access. The main creator of this dataset is Reddit user OminousLatinWord. For convenience, I converted the dataset to Parquet files and uploaded it to Hugging Face. Source Data & Attribution Creator: Created and originally open-sourced by Reddit u
Hugging Face Datasets2026 · Table · Parquet
TikTok Videos, 4.5 BillionMirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies. TikTok Videos: 4.5 billion posts with engagement metrics 4.5 billion TikTok video records with captions, engagement counts, sound identifiers and timing. Collected from TikTok's mobile API over roughly three weeks. Every content_id appears exactly once. This is the
Hugging Face Datasets2026 · Table · Parquet
Infoton Identity — 735 Human ProteinsAn Infoton Dose of RxRx — 735 Human Proteins Version: 1.0.0 Released: 2026-09-06 Author: Walker, January Natatia — Infoton DOI: 10.5281/zenodo.18210355 ORCID: 0009-0000-6843-2051 Overview The Infoton Identity Physics Engine computes in seconds per protein and aligns with what cellular imaging measures. The coordinates in the An Infoton Dose of RxRx dataset were derived before correlation was run w
Hugging Face Datasets2026 · Table · Parquet
SIGS Symbolic Expression Latent CorpusSIGS symbolic expression latent corpus This dataset contains 23,695 grammar-generated symbolic expressions used by the SIGS Grammar-VAE, together with their 32-dimensional latent statistics and variable-based mathematical classes. Dataset structure Each row contains: id: stable row index; expression: symbolic expression generated by the SIGS grammar; math_class: one of CONSTANT, TEMPORAL_1D, SPATI
Hugging Face Datasets2026 · Table · Parquet
LITCOIN Proof-of-Research CorpusLITCOIN Proof-of-Research Corpus 191,484,662 AI research submissions, produced by 81,224 anonymous contributors and 470 model variants competing against each other, every row executed in a sandbox and scored. This is the complete output of the LITCOIN protocol, which ran on Base from March to August 2026. Autonomous AI agents were paid in a permissionless token to solve real optimization problems
Hugging Face Datasets2026 · Table · Parquet
Autonomous AI Agents & Multi-Agent Swarms Dataset (2023-2026)🤖 Autonomous AI Agents & Multi-Agent Swarms Dataset (2023–2026) Sample dataset of 30 audit-verified research papers covering Autonomous AI Agents, Multi-Agent Swarms, Tool Calling, and Model Context Protocols (MCP) with 384d PyTorch embeddings. 🛒 Full 1,000 Paper B2B Dataset Available on Gumroad Get the complete 3-year dataset (1,000 papers + VRAM & Execution Modes + SQLite/CSV/Parquet + Quickstar
Hugging Face Datasets2026 · Image
tomato_leavesTomato Leaves Dataset Overview This dataset contains images of tomato leaves categorized into different classes based on the type of disease or health condition. The dataset is divided into training, validation, and test sets, with a ratio of 8:1:1. The classes include various diseases as well as healthy leaves. The dataset includes both augmented and non-augmented images. Dataset Structure The da
Hugging Face Datasets2026 · Table · CSV
SynPerFormSynPerForm SynPerForm is a paired Persian dataset for formality style transfer. Each informal text is paired with a freely written formal rewrite that preserves its meaning without requiring lexical or structural equivalence. The formal rewrites were generated with OpenAI GPT-5.6 Luna. Columns Informal: the original informal Persian text. Formal: a free formal rewrite that preserves the original m
Hugging Face Datasets2026 · Table · Parquet
Common Crawl Web Graph EmbeddingsWeb Graph Embedings from Common Crawl's Host-level Hyperlink Graph Dense 128-dimensional embeddings for 52,913,544 web hosts, learned by link prediction on the Common Crawl host-level hyperlink graph release cc-main-2025-26-nov-dec-jan. Vectors are L2-normalized and served in float16; similarity is cosine (a dot product on the unit vectors). The dataset contains hosts with link degree >= 8 (total
Hugging Face Datasets2026 · Image
IMF Technical Assistance Reports — Recommendation Process CorpusIMF Technical Assistance Reports — Recommendation Process Corpus A page-grounded research corpus of 780 IMF technical-assistance report records. It contains source PDFs, layout-aware Markdown, page-level text, extracted visuals, metadata, observations, recommendations, and labeled links between observations and recommendations. Required acknowledgement All research, publications, datasets, models,
Hugging Face Datasets2026 · dataset
Comprehensive DDoS Datasets based on Feature ExtractionDataset Card for Comprehensive_Feature_Extraction_DDoS_Datasets This dataset card aims to be provided preprocessed five published DDoS datasets based on three feature extracted methods, including correlation, IM, and UFS. Dataset Description The imperative for robust detection mechanisms has grown in the face of increasingly sophisticated Distributed Denial of Service (DDoS) attacks. This paper in
Hugging Face Datasets2026 · dataset
Rust StackOverflow Vector DatasetRust StackOverflow Vector Dataset Summary This Hugging Face dataset repository contains the Rust shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifacts are intended for
Hugging Face Datasets2026 · Table · Parquet
MINDSETMINDSET MINDSET is the pretraining dataset for MIND, a coordinate-only location encoder distilled from static location encoder teachers and annual AlphaEarth Foundations (AEF) embeddings. We release the embeddings at the 12.1M training coordinates. The dataset contains 12,099,072 land coordinates in WGS84. Coordinates are dense around cities and not uniformly sampled over land. The files are in Ge
Hugging Face Datasets2026 · Text · gated
PII Masking, Health & Medical Information (PHI) with Asia Pacific👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Health & Medical Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations La
Hugging Face Datasets2026 · Table · CSV
uw-math-ai/math-graphMath-Graph Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency graph spanning both informal and formal mathematics. On the informal side it parses millions of theorem-like environments from mathematics arXiv and recovers directed dependency edges within and across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed declaration