Hugging Face Datasets2026 · Image
PolyTopoBenchPolyTopoBench PolyTopoBench is a benchmark for complex vector polygon generation from remote-sensing imagery. It evaluates whether a model can produce complete polygon topology, including interior rings (holes), rather than only exterior boundaries. Paper: PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery, NeurIPS 2026 (Evaluations and Datasets Track) Cod
Hugging Face Datasets2026 · Table · Parquet
TMB-S3TMB-S3: Temporal Modelling for Burn Scars on Sentinel-3 TMB-S3 is a burned-area segmentation dataset of 246 European wildfire activations (2016–2025) of the Copernicus Emergency Management Service (CEMS), split into 770 bounding boxes of about 25 × 25 km. Each bounding box contains a Sentinel-3 OLCI time series (one pre-fire acquisition followed by the post-fire acquisitions), the CEMS burned-area
Hugging Face Datasets2026 · Text
Indian Railways stop-level running and delay historyIndian Railways: stop-level running and delay history One row is one train, at one station, on one day. What time it was due, what time it actually turned up, and how late that was. One CSV per train, named by its five digit train number. 12624.csv holds every recorded run of train 12624, oldest first. Every file has the same 24 columns. The files are split across two folders, because the Hub allo
Hugging Face Datasets2026 · Table · Parquet
Major-TOM/Core-S1RTC-UniverSatGlobal coverage of the embeddings, coloured by the top-3 principal components of the 768-d UniverSat vectors (mapped to RGB). Distinct colours mark distinct backscatter regimes — dense tropical forest (magenta), arid and desert surfaces (green), boreal and temperate land (tan), open water and other smooth low-backscatter surfaces (blue). Core-S1RTC-UniverSat 🛰️ Dataset Modality Number of Embedding
Hugging Face Datasets2026 · dataset
MykolaL/FBIS-73MFBIS-73M Field Boundary Instance Segmentation - 73M dataset (FBIS-73M) - large-scale, multi-resolution dataset comprising 1,478,096 high-resolution satellite image patches (0.25 m – 10 m) and 74,015,889 instance masks of individual fields. Dataset Structure The dataset has three splits: train, test, and a 100-country zero-shot test set. Each split is described by a list file, and the actual image/
Hugging Face Datasets2026 · Text
justuskarlsson/FireCompFireComp: Next-Day Fire Spread Global benchmark for next-day wildfire spread prediction from 375 m VIIRS active-fire detections, with ERA5 weather, GFS forecasts and Alpha Earth terrain embeddings. 256×256 patches, 9 regions, 2017–2025. Code, documentation, loaders and paper: https://github.com/justuskarlsson/FireComp Paper This dataset accompanies the FireComp paper, accepted and presented at GCP
Hugging Face Datasets2026 · Image
miti360Miti360: An integrated dataset combining remote sensing, ground measurements and weather data for improved reforestation monitoring Introduction In the era of artificial intelligence, machine learning combined with remote sensing and ground measurements offers unprecedented opportunities to enhance forest monitoring through faster, more accurate biomass estimation and individual tree analysis. Des
Hugging Face Datasets2026 · dataset
Plumbing Permit Dataset - Multi-City Building Systems Records - EmbedEarthPlumbing Permit Dataset - Multi-City Building Systems Records - EmbedEarth A geospatial open-data release of 223,000 plumbing permit records from 2018, 2019, 2020, 2021, 2022, 2023. This dataset was created and distributed by EmbedEarth, programmable geographic infrastructure for searching, retrieving, and computing across the physical world. The release contains records of permitted plumbing work
Hugging Face Datasets2026 · Image
AI-TOD-v2AI-TOD-v2 AI-TOD-v2, the tiny-object detection benchmark in aerial images, packed once with the official v2 annotations kept whole, so it loads in one line and no data path has to be configured: from datasets import load_dataset ds = load_dataset("shijli/aitod-v2") # 11214 train / 2804 validation / 14018 test AI-TOD cuts 28036 images of 800 x 800 pixels from xView, DOTA-v1.5, VisDrone2018-Det, Air
Hugging Face Datasets2026 · Table · Parquet
michiyomi — Tokyo streetscape verbalization open data / 東京都 街路言語化オープンデータmichiyomi — Tokyo streetscape verbalization open data English Overview michiyomi pairs coordinates with structured Japanese descriptions of physical streetscapes visible in public Mapillary imagery. A vision-language model (VLM) verbalized only what is visible in each image: no map, address, place name, facility name, statistics, or other external knowledge was injected. Release 2026-09-13-r1 cont
Hugging Face Datasets2026 · Image
FZI-AURAFZI-AURA A multimodal autonomous-driving dataset featuring the largest LiDAR sensor suite of any public autonomous-driving dataset. Download | Python SDK | Data format | License 2,473 scenes | 13.70 hours | 8 cameras | up to 12 LiDARs | 4.11 million 3D-box annotations | 30.01 billion segmented LiDAR points Public preview release: 1,081 out of the 2,473 FZI-AURA scenes are currently available. The
Hugging Face Datasets2026 · Text
US Parcel Layer — The Landrecords.us Nationwide Parcel DatasetUS Parcel Layer — The Landrecords.us Nationwide Parcel Dataset The Landrecords.us National Parcel Dataset is a comprehensive, standardized geospatial dataset aggregating ~157 million parcel boundaries and their associated land-ownership and taxation attributes, harmonized from thousands of local jurisdictions across the United States into a single national schema. Each record is a land parcel: a p
Hugging Face Datasets2026 · Image
taylor-geospatial/CoordBenchCoordBench A unified benchmark suite for evaluating location encoders such as SatCLIP, GeoCLIP, Climplicit, and MIND. The dataset contains 40 normalized source tables from 13 source families. The paper's evaluation suite uses 52 datasets and 78 prediction targets drawn from this mirror. The source files previously lived across GitHub, figshare, GCS, Socrata, Zenodo, and Google Drive. Intended use
Hugging Face Datasets2026 · Image · gated
SkyLumeSkyLume SkyLume is a large-scale real-world UAV dataset for urban scene reconstruction under varying illumination. It captures the same urban regions across morning, noon, and afternoon flights, pairing high-resolution five-direction UAV imagery with LiDAR-derived geometry for robust 3D reconstruction, novel view synthesis, inverse rendering, and cross-time consistency research. Project page: http
Hugging Face Datasets2026 · Text
CMA-LocCMA-Loc This repository stores the CMA-Loc dataset as uncompressed tar shards plus the original JSON annotation files. Layout data/*.tar # image shards, paths inside each tar keep the original urban/<city>/<modality>/... layout metadata/*.json # original annotation split files manifests/shards.jsonl # one row per tar shard manifests/files.jsonl # one row per image file with its shard mapping… See
Hugging Face Datasets2026 · Table · Parquet
MINDSETMINDSET MINDSET is the pretraining dataset for MIND, a coordinate-only location encoder distilled from static location encoder teachers and annual AlphaEarth Foundations (AEF) embeddings. We release the embeddings at the 12.1M training coordinates. The dataset contains 12,099,072 land coordinates in WGS84. Coordinates are dense around cities and not uniformly sampled over land. The files are in Ge
Hugging Face Datasets2026 · Image
GroundSetGroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data GroundSet is a large-scale Earth Observation dataset grounded in verifiable cadastral vector data, designed to bridge the gap in fine-grained spatial understanding for modern Multimodal Models. The dataset is built upon high-resolution (20 cm) optical aerial orthophotos and legally verified vector data provided by t
Hugging Face Datasets2025 · dataset
zhu-xlab/So2Sat-LCZ42So2Sat-LCZ42 A copy of the Fourth version of the So2Sat-LCZ42 dataset, pairing the second version (culture-10) with geolocation: Training: 42 cities around the world Validation: western half of 10 other cities covering 10 cultural zones Testing: eastern half of the 10 other cities Description of the files First extract [split].h5.gz to [split].h5 with gunzip [split].h5.gz. training.h5: training da
Hugging Face Datasets2025 · dataset · gated
IWI-Africa Multimodal DatasetNOTICE: By requesting access to this dataset, you are affirming and certifying that you have obtained the necessary approval from The Demographic and Health Surveys (DHS) Program to access and use the underlying DHS microdata upon which this dataset is based. Such approval requires registration as a user on The DHS Program's official website, submission of a detailed research project description f
Hugging Face Datasets2025 · dataset
DASPDataset Card for DASP Dataset Description The DASP (Distributed Analysis of Sentinel-2 Pixels) dataset consists of cloud-free satellite images captured by Sentinel-2 satellites. Each image represents the most recent, non-partial, and cloudless capture from over 30 million Sentinel-2 images in every band. The dataset provides a near-complete cloudless view of Earth's surface, ideal for various geos
Hugging Face Datasets2025 · Image
BrightOverview BRIGHT is the first open-access, globally distributed, event-diverse multimodal dataset specifically curated to support AI-based disaster response. It covers five types of natural disasters and two types of man-made disasters across 14 regions worldwide, with a particular focus on developing countries. About 4,200 paired optical and SAR images containing over 380,000 building instances in
Hugging Face Datasets2024 · Image
Major-TOM/Core-S2L1CCore-S2L1C Contains a global coverage of Sentinel-2 (Level 1C) patches, each of size 1,068 x 1,068 pixels. Source Sensing Type Number of Patches Patch Size Total Pixels Sentinel-2 Level-1C Optical Multispectral 2,245,886 1,068x1,068 2.56 Trillion Content Column Details Resolution B01 Coastal aerosol, 442.7 nm (S2A), 442.3 nm (S2B) 60m B02 Blue, 492.4 nm (S2A), 492.1 nm (S2B) 10m B03 Green, 559.8 n
Hugging Face Datasets2024 · Image
Major-TOM/Core-S2L2ACore-S2L2A Contains a global coverage of Sentinel-2 (Level 2A) patches, each of size 1,068 x 1,068 pixels. Source Sensing Type Number of Patches Patch Size Total Pixels Sentinel-2 Level-2A Optical Multispectral 2,245,886 1,068 x 1,068 (10 m) > 2.564 Trillion Content Column Details Resolution B01 Coastal aerosol, 442.7 nm (S2A), 442.3 nm (S2B) 60m B02 Blue, 492.4 nm (S2A), 492.1 nm (S2B) 10m B03 Gr
Hugging Face Datasets2023 · Image
LEVIR CD+LEVIR CD+ The LEVIR-CD+ dataset is an urban building change detection dataset that focuses on RGB image pairs extracted from Google Earth. This dataset consists of a total of 985 image pairs, each with a resolution of 1024x1024 pixels and a spatial resolution of 0.5 meters per pixel. The dataset includes building and land use change masks for 20 different regions in Texas, spanning the years 2002