{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:video-classification",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Text
Synthetic Street ScenesSynthetic Street Scenes Mehmet Kerem Turkcan Columbia University, Center for Smart Streetscapes (CS3) Synthetic Street Scenes is a collection of 22 video datasets of street situations that are rare, dangerous or impractical to film: street flooding, snow cover, traffic jams, collisions and near misses, storm damage, blocked bike lanes, tampering with roadside sensors, and civic maintenance problem
Hugging Face Datasets2026 · Table · Parquet
XPlanner BenchmarkXPlanner-Benchmark XPlanner-Benchmark is the portable release of the X-Planner 1,500-episode evaluation benchmark. It contains synchronized multi-view robot-manipulation videos and the episode-level task, subtask, action, scene, duration, and complexity metadata used by X-Planner. Contents 1,500 episodes 3,490 MP4 video references 167 source dataset identifiers 525 unique task names 31 task classe
Hugging Face Datasets2026 · Image
LSI-108KLSI-108K Interaction-derived supervision for spatial reasoning LSI-108K is the dataset introduced in Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World. It contains 107,518 verifiable QA pairs organized as a three-level spatial interaction curriculum. Overview Presentation A three-level curriculum Split Samples Curriculum role L1 15,109 Passive… S
Hugging Face Datasets2026 · Text
Eidon Tracker POVEidon Tracker POV 1,274 hours of egocentric video paired with 7-point IMU arm tracking, recorded during ordinary household work. Contributors wore a head-mounted camera and a seven-sensor IMU harness while doing real chores in their own homes: laundry, cleaning, dishes, cooking. Each recording pairs first-person video with 24 Hz orientation data for both hands, both forearms, both upper arms, and
Hugging Face Datasets2026 · Image
RekaAI/RekaDaily-10k-processedRekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second per
Hugging Face Datasets2026 · dataset
WGO-BenchWGO-Bench MacroData's WGO-Bench is a benchmark for evaluating timestamped subtask annotations from robot and egocentric manipulation videos. MacroData describes the benchmark and its annotation approach in Segmenting Robot Video into Actionable Subtasks. The upstream repository's summary.json states that its timeline reannotations were applied only to the 25 homer_* episodes, while the DROID and G
Hugging Face Datasets2026 · dataset · gated
SoccerTrack v2SoccerTrack v2 Ten university-level soccer matches (934 minutes) recorded by fixed panoramic camera systems whose field of view spans the entire pitch, released together with frame-level game state annotations and player-linked ball action events on the same footage. Contents Folder Contents videos/ 20 panoramic half-match videos (two 45-minute periods per match) gsr/ Game state annotations: one J
Hugging Face Datasets2026 · Image
RoboPulse++RoboPulse++ RoboPulse++ is an interval-level benchmark introduced in PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment for evaluating progress judge models throughout complete robot manipulation trajectories. This Hugging Face release contains 700 episodes with natural-language task instructions, temporally ordered observations, and human-annotated progress intervals. Overview RoboPulse++
Hugging Face Datasets2026 · dataset
PhysicalAI-Event-VideosPhysicalAI-Event-Videos PhysicalAI-Event-Videos is a video-anomaly and event dataset for developing text-to-video anomaly-search and safety-event-understanding systems. Version 1.0 provides structured annotations for 1,612 parent-video records and 3,796 labeled chunks, with 71,023 captions and queries across person-attribute, general-caption, anomaly, and action search tasks. The redistributable m
Hugging Face Datasets2026 · dataset · gated
PersoMoni DatasetPersoMoni Dataset A Comprehensive Video-Based Benchmark Dataset for Fine-grained Personality Assessment with 15 Trait Dimensions Language: English | 中文 How to apply The application form is displayed in the Hugging Face access panel at the top of this dataset page. Sign in with the applicant's personal Hugging Face account, complete all identity, affiliation, research, ethics, and data-security fie
Hugging Face Datasets2026 · dataset · gated
HarassGuard: Social VR Harassment Behavior Video DatasetHarassGuard Dataset Social VR Harassment Behavior Video Dataset The HarassGuard Dataset is a vision-based video dataset developed for detecting physical harassment behaviors in Social Virtual Reality (Social VR) environments using visual information. The dataset was collected through the user study conducted for the following paper: HarassGuard: Detecting Harassment Behaviors in Social Virtual Rea
Hugging Face Datasets2026 · Table · CSV
Real-TurnTurkReal-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attr
Hugging Face Datasets2026 · Table · Parquet
Factory Manipulation VideosFactory manipulation videos Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks. Contents Seven continuous takes, 109 minutes in total. Task Station Worker Duration File cardboard manipulation 01 041 23.6 min cardboard_manipulation_station01_worker041.mp4 cardboard manipulation 04 026 16.5 min cardbo
Hugging Face Datasets2026 · dataset · gated
EgoSuite-Open100K - EgoDemoEgoDemo A 50-hour sample from EgoSuite-Open100K, covering every annotated subset plus two raw-video variants. Collection · EgoStandard · EgoPro · Project page Explore EgoSuite-Open100K ↗ EgoSuite-Open100K Overview Collection: EgoSuite-Open100K SKU Sub-SKU Format Planned Duration EgoStandard EgoStand Hand Pose 80,000 h… See the full description on the dataset page: https://huggingface.co/datasets/L
Hugging Face Datasets2026 · dataset · gated
EgoSuite-Open100K - EgoStandardEgoStandard The 90,000-hour head-view line of EgoSuite-Open100K. Data Bucket · Collection · EgoDemo · EgoPro · Project page Explore EgoSuite-Open100K ↗ Data location: EgoStandard is distributed through the LightwheelAI/EgoStandard Bucket. This Git repository is the dataset card and access point; download the data from the Bucket. Overview EgoStandard pairs head-view egocentric video with synchroni
Hugging Face Datasets2026 · dataset · gated
EgoSuite-Open100K - EgoProEgoPro The 10,000-hour head-and-wrist line of EgoSuite-Open100K. Data Bucket · Collection · EgoDemo · EgoStandard · Project page Explore EgoSuite-Open100K ↗ Data location: EgoPro is distributed through the LightwheelAI/EgoPro Bucket. This Git repository is the dataset card and access point; download the data from the Bucket. Overview EgoPro pairs synchronized head- and wrist-view video with 3D han
Hugging Face Datasets2026 · dataset
ACE-Data-0ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore · ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video · Full story, figures, and interactive examples on the blog What this is Learning to act in the physical… See the
Hugging Face Datasets2026 · Text
SurgAtlasSurgAtlas SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery [Paper] SurgAtlas is a large-scale surgical video-language dataset built from publicly available surgical videos on YouTube. It contains 15,291 videos and 2,391 hours of surgery, spanning 18 surgical specialties and more than 5,000 procedure types. SurgAtlas includes 6,182 ope
Hugging Face Datasets2026 · Text · gated
LRS3-TED (verified mirror)LRS3-TED — verified mirror A mirror of the LRS3-TED dataset (Lip Reading Sentences 3), preserved because the official distribution has been discontinued. This repository adds no new data: it is a re-hosted copy with a full verification report against the official file list, so you know exactly what is and is not here. Attribution LRS3-TED was created by Triantafyllos Afouras, Joon Son Chung and An
Hugging Face Datasets2026 · Image · gated
SAFER-ActivitiesSAFER-Activities A Dataset for Smart Assessment of Fall Events and Routine Activities (ECCV 2026). SAFER-Activities is a dataset for fall detection and physical activity monitoring from video, with a dedicated subset for wheelchair users. It provides frame-level action annotations (precise start/end of every action) over long, untrimmed, multi-camera recordings. File structure . ├── raw/ │ ├── Saf
Hugging Face Datasets2026 · dataset
Children Gait VideoDecoding Children's Gait Behavior ECCV 2026 Yifan Shen1,2,*, Boyi Li1,*, Meihuan Huang2,3,4,*, Yuanzhe Liu1,*, Xu Cao1,2,*,§, Jinyang Jin1, Zhengyuan Li1, Anglin Liu5, Junho Kim1, Jingyuan Zhu2, Fangzhou Lan2, Jianguo Cao2,3, Jintai Chen5, Ismini Lourentzou1, James M. Rehg1,† 1 University of Illinois Urbana-Champaign 2 PediaMed AI 3 Shenzhen Children's Hospital 4 Hong Kong Polytechni
Hugging Face Datasets2026 · Image
RoboJudge DataRoboJudge Data This repository hosts the video assets and canonical metadata for evaluating Physical Adherence (PA) and Instruction Alignment (IA) in generated embodied-manipulation videos. Current RoboJudge release Use robojudge_release/ for the paper release: Path Contents robojudge_release/train/physical_adherence.json 12,351 PA training records robojudge_release/train/instruction_alignment.jso
Hugging Face Datasets2026 · Image
TikTok Creative Center — Top Ads (v0.2.0, full corpus)TikTok Creative Center — Top Ads (v0.2.0, full corpus) Full corpus scraped from the TikTok Creative Center Top Ads surface (country=US, period=180, sort=for_you). One row per ad_id with embedded video bytes (mp4), cover image (jpg/png), per-second engagement curves, and TikTok-reported metadata. 61,789 ads. Sister dataset liangyuch/ttcc-v0_1_0 is the legacy 935-ad smoke set with the same schema —
Hugging Face Datasets2026 · Image · gated
IndexTeam/CASTER-BenchCASTER-Bench CASTER-Bench is a human-annotated multimodal benchmark for Community-Aware Assessment of Social Textual Engagement and Resonance (CASTER) — a task that evaluates whether User-Generated Content (UGC) achieves positive community resonance, going beyond traditional aesthetic-focused Video Quality Assessment (VQA). This benchmark is introduced in the ACL 2026 paper: "Community-Aware Asses
Hugging Face Datasets2026 · Image · gated
UII-AI/MedVidU_ECCV2026_TrainValECCV 2026 Workshop on Medical Video Understanding (MedVidU @ ECCV 2026) — Train / Val Split This is the public train / val split for the MedVidU Challenge at the ECCV 2026 Workshop on Medical Video Understanding. This split is derived from the benchmark introduced in MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding (CVPR 2026). Participants are free to use t