{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:image-feature-extraction",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Image · gated
Car Damage ImagesCar Damage Images A raw image collection for vehicle damage assessment. Unlabeled: these images have no annotations yet and are intended as source material for labeling or pre-training. Structure images/ 001/ images_001.jpg images_002.jpg ... 002/ ... ... 683 folders thumbnails/ 001/ thumbnail_001.jpg thumbnail_002.jpg ... ... 683 folders Branch Files Folders Size… See the full description on the
Hugging Face Datasets2026 · Table · Parquet
PubMed-OphthaPubMed-Ophtha PubMed-Ophtha is a hierarchical ophthalmic vision-language dataset built from open-access articles in PubMed Central. Figures are extracted directly from the article PDFs at full resolution and decomposed into panels, panel identifiers, and individual images; figure captions are split into panel-level subcaptions. The result is a panel-centric corpus of 102,023 panels from 35,544 fig
Hugging Face Datasets2026 · Image
MONETDataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domai
Hugging Face Datasets2026 · dataset
African Medical Records (AMR): Nigerian Handwritten Clinical RecordsWhy AMR Exists Clinical documentation across much of Africa is still handwritten, and the world's OCR and handwritten-text-recognition (HTR) systems have almost never seen it. Models trained on Western clinical forms or clean printed text fail on real Nigerian ward notes, prescriptions, and observation charts, where handwriting styles, abbreviations, drug names, and document formats differ sharply
Hugging Face Datasets2026 · Image · gated
OpenTMEOpenTME: Open-Access Tumor Microenvironment Profiles from TCGA OpenTME is an open-access project by Aignostics for academic researchers. It provides comprehensive spatial outputs for whole slide images (WSIs) of H&E-stained, formalin-fixed, paraffin-embedded slides from The Cancer Genome Atlas (TCGA). OpenTME is powered by Atlas H&E-TME – a computational pathology application developed by Aignosti
Hugging Face Datasets2026 · Image
TCGA WSI UNI2H FeaturesTCGA WSI UNI2H Features Dataset Summary This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research. Data is organized by project (for example TCGA-HNSC) and currently exposes: features/ containing H5 feature files with tile-level embeddings vis/ containing overlay images for
Hugging Face Datasets2026 · Image · gated
FOMO300KFOMO300K: Brain MRI Dataset for Large-Scale Self-Supervised Learning with Clinical Data Dataset paper preprint: A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning. https://arxiv.org/pdf/2506.14432. Updates April 3, 2026: V1.1 released. MGH-Wild has changed license and is no longer part of FOMO300K. The new dataset is now available and totals 306,20
Hugging Face Datasets2026 · Image
Synthetic ResistorsIf you want a larger dataset, or want to change some parameters see: https://github.com/tylerebowers/synthetic_resistor_generation
Hugging Face Datasets2025 · Image
AgentDS-Manufacturing🏭 AgentDS-Manufacturing This dataset is part of the AgentDS Benchmark — a multi-domain benchmark for evaluating human-AI collaboration in real-world, domain-specific data science. AgentDS-Manufacturing includes sensor and operational data for 3 challenges: Predictive maintenance (equipment failure within 24h) Production delay forecasting Quality cost estimation 👉 Files are organized in the Manufac
Hugging Face Datasets2025 · dataset · gated
vigil1917/GazeGeneGazeGene: Large-scale Synthetic Gaze Dataset with 3D Eyeball Annotations Yiwei Bao, Zhiming Wang, Feng Lu {baoyiwei, zy2306418, lufeng}@buaa.edu.cn This work is presented by Phi-A2I Lab. We propose the GazeGene dataset: a large-scale synthetic gaze estimation dataset providing over 1M images with 3D eyeball annotations and accurate gaze annotations. Abstract Thanks to the introduction of large-sca
Hugging Face Datasets2025 · Image
NVIDIA Physical AI SimReady Warehouse OpenUDS DatasetNVIDIA Physical AI SimReady Warehouse OpenUSD Dataset Dataset Version: 1.1.0 Date: May 18, 2025 Author: NVIDIA, Corporation License: CC-BY-4.0 (Creative Commons Attribution 4.0 International) Contents This dataset includes the following: This README file A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asse
Hugging Face Datasets2024 · Image · gated
LouisChen15/ConstructionSiteDataset Card for ConstructionSite 10k Dataset summary The dataset consists of a total of 10,013 construction site images and their annotations. Among them, 7,009 images are assigned to the training split while 3,004 images are assigned to the test split. If you use this dataset, we would appreciate you citing our work. See Citation information. We developed the dataset to test how well vision lang
Hugging Face Datasets2024 · dataset · gated
Amber-River/Pixiv-2.6MPixiv 2.6M specific This dataset contain 2.6M images from Pixiv. Introduction This dataset aims to add some concepts which is lacked or banned in danbooru. Images Content The images is collected with search cralwer made with selenium and python I use some tags which may lack in danbooru for example "scenery" or some tags for "explicit contents" to do the search and collect the url of result than d