{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:audio-classification",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Table · Parquet
SANJAYKISHORE/TORGO-databaseThe TORGO Database: Acoustic and articulatory speech from speakers with dysarthria Dataset Summary This database only includes the short words and restricted sentence portion of the TORGO dataset. For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html. Transcripts have been normalized to remove punctua
Hugging Face Datasets2026 · Text
Star Wars dialogue annotations (master cue table)Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with
Hugging Face Datasets2026 · Table · Parquet
DeepASMR-NSpeechDeepASMR-NSpeech A Fine-Grained Benchmark for Non-Speech ASMR Generation 73,829 ten-second clips · 202.6 hours · 37 fine-grained actions 🎧 Interactive Demo · ⌘ Code · 📄 Paper: coming soon Dataset summary DeepASMR-NSpeech is a non-speech ASMR audio dataset with structured Subject-Verb-Object annotations and the SVO-AQA audio question-answering benchmark. All clips are 10 seconds long and cover 37 f
Hugging Face Datasets2026 · Text
Voice Isolation Benchmark - Processed AudiosVoice Isolation Benchmark – Processed Audios Audio samples processed by four Krisp Voice Isolation models. Each scenario folder contains subfolders for every model, with one processed .wav file per original sample. Voice Isolation Models VI 2.5 Default (vi_2_5_default) The main Voice Isolation model. Strongest at removing noise and other speakers. Goes fully silent when only a bystander is talking
Hugging Face Datasets2026 · Text · gated
Basis ConversationsDataset Card for Basis Conversations 1500 Listen first: sample conversations Overview Basis Conversations 1500 is a multi-party, multilingual, full duplex conversational speech dataset. Each conversation includes up to 4 simultaneous speakers, each with channel-separated, 48 kHz audio. The median conversation lasts 33 minutes and 2,645 unique speakers are represented. Multi-party: a conversation s
Hugging Face Datasets2026 · Table · Parquet
InteractSpeech English Text and TimelineInteractSpeech: English Text-and-Timeline Release This is the English text-and-timeline release associated with InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (Findings of EMNLP 2025). InteractSpeech is designed for real-time spoken-dialogue interaction, including interruptions, backchannels, pauses, gaps, overlaps, and turn transitions. The paper describes an appro
Hugging Face Datasets2026 · Table · Parquet
URGENT 2026 Speech Quality AssessmentURGENT 2026 Speech Quality Assessment The subjective listening-test dataset from ICASSP 2026 URGENT Track 1: Universal Speech Enhancement, containing 5,040 ACR samples, 12,600 CCR comparisons, and 141,386 individual ratings across 840 utterances, nine languages, and six systems. Each scored sample includes the audio, its MOS or CMOS, individual listener scores, and utterance metadata. Load the dat
Hugging Face Datasets2026 · Table · Parquet
BEAT Audio-to-Blendshape (log-mel / ARKit 51)BEAT Audio-to-Blendshape Paired 80-bin log-mel speech features and 51-dim ARKit blendshape tracks, derived from BEAT v1 (english v0.2.1). 6,357 speech segments / 55.5 hours / 30 speakers. This is the exact data our audio-driven face model was trained on, in the exact train/validation partition it used — not a re-cut sample. Contents Segments 6,357 (train 5,970 · validation 387) Duration 55.5 h (tr
Hugging Face Datasets2026 · dataset · gated
ChaashiniChaashini (चाशनी) Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. Total: 1,391,986 clips · 2955.25 hours · 33 langu
Hugging Face Datasets2026 · Table · Parquet
VeriSpeakVeriSpeak VeriSpeak is a spoken-statement factual-verification benchmark. Each example is a short synthesized speech clip of a single declarative sentence about a public figure, labeled correct or incorrect depending on whether the spoken statement is factually true. The task: given the audio (and optionally its transcript), decide whether the claim it makes is accurate. It targets speech-native f
Hugging Face Datasets2026 · Table · CSV
Real-TurnTurkReal-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attr
Hugging Face Datasets2026 · Text
Beyond Binary Instrument QA🎵 Beyond Binary Instrument QA:Probing Instrument Grounding in Music Audio-Language Models Yujun Lee · Joonhyeok Shin · Hyoeun Kim · Kyuhong Shim Sungkyunkwan University 📄 arXiv | 🤗 Dataset Benchmark release. Five complementary evaluation configurations test instrument presence, reduced genre-prior reliance, fine-grained discrimination, long-context multi-label recognition,
Hugging Face Datasets2026 · Text
AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts)AppTek Call-Center Dialogues — Travel and Hospitality (No Transcripts) This is a filtered derivative of AppTek Call-Center Dialogues, prepared for a specific use case. Changes from the source dataset Restricted the dataset to the travel and hospitality domains. Removed the transcript field (text) entirely. Kept the original audio and the domain, gender, and accent metadata. Preserved the source da
Hugging Face Datasets2026 · Table · Parquet
YO-CPT-kkYO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities,
Hugging Face Datasets2026 · Table · Parquet
MAJEPPA: Piano Performance MIDI with Score Pairings and Expertise LabelsMAJEPPA 4,449 solo piano performance MIDI files (279 hours, ~5.0M notes) transcribed from amateur and professional piano recordings, each labelled with the performer's expertise level and recording context. 4,207 (95%) paired with a reference score in MIDI — 886 scores covering 880 distinct works and movements by 52 composers 3,862 (87%) carry a DTW score↔performance alignment — 20 million timesta
Hugging Face Datasets2026 · dataset
ACE-Data-0ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore · ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video · Full story, figures, and interactive examples on the blog What this is Learning to act in the physical… See the
Hugging Face Datasets2026 · Table · Parquet
Zenless VoiceZenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename
Hugging Face Datasets2026 · Text
Multilingual Indian Conversational SpeechMultilingual Indian Conversational Speech A dataset of naturalistic, spontaneous two-speaker conversations across 13 Indian languages, with segment-level transcripts, speaker profiles, timestamps, and recording metadata. Designed for ASR, TTS, speaker diarization, and conversational speech research. Languages (13) Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Odia, Punja
Hugging Face Datasets2026 · dataset
InterPet4DInterPet4D (v1) Authors: Yichen Peng*, Jyun-Ting Song*, Chen-Chieh Liao*, Kris Kitani, Hideki Koike, Erwin Wu *Equal contribution. InterPet4D is a multimodal, ego-centric dataset of natural human–pet (dog) interactions. Each clip provides time-synchronized audio, SMPL human body motion, MANO hand motion, pet skeletal motion, and SMAL pet body parameters, enabling research on cross-species interact
Hugging Face Datasets2026 · dataset
Multilingual-NLP/YUE-PUB-SpeechYUE-PUB-Speech This repo contains the data for our paper YUE-PUB-Speech: A Speech-based Pragmatic Understanding Benchmark for Cantonese (Interspeech 2026). Links 🤗 Dataset: https://huggingface.co/datasets/Multilingual-NLP/YUE-PUB-Speech 💻 Code: https://github.com/swaggy66/Yue-PUB-Speech The benchmark consists of 14 tasks collected from 7 existing pragmatic datasets, covering different forms of pra
Hugging Face Datasets2026 · Image
dolphinteam/OpenWhistle-CNNOpenWhistle CNN Dataset dolphinteam/OpenWhistle-CNN is the public CNN dataset used for binary dolphin whistle detection. It contains audio windows, spectrogram images, and binary labels: noise (label=0) whistle (label=1) The main dataset is the complete session-disjoint dataset used for training and evaluation. A smaller deterministic review-sample config is also provided so reviewers can inspect
Hugging Face Datasets2026 · Image
dolphinteam/OpenWhistle-Classification-FinetuningOpenWhistle Classification Finetuning Dataset dolphinteam/OpenWhistle-Classification-Finetuning is the public classification finetuning dataset used for dolphin whistle identity classification. It contains short whistle clips, whistle-level metadata, fundamental-frequency tracks, rendered F0 spectrograms, and integer class labels. The main reviewer-facing subset is the balanced balanced config. It
Hugging Face Datasets2026 · Text
Audio2Tool — Spoken Tool-Calling BenchmarkAudio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dat
Hugging Face Datasets2026 · Text
WorldSpeechWorldSpeech 🎉 WorldSpeech has been accepted to NeurIPS 2026! 🎉See the paper on arXiv. A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an
Hugging Face Datasets2026 · Text · gated
Hinglish Code-Switched Conversational Dataset v1Hinglish Code-Switched Conversational Dataset v1 Overview This dataset contains structured Hinglish conversational voice data built to reflect how people actually speak in real-world interactions. Most speech datasets are clean, scripted, or heavily processed. That works in controlled testing, but it breaks in production where speakers interrupt each other, switch languages, use regional accents,