{
"version": "1.0",
"query": {
"text": null,
"operator": "and",
"filters": [
{
"field": "concept",
"op": "descendant_of",
"value": "hf_task:automatic-speech-recognition",
"evidence": null
}
],
"evidence_policy": "standard"
},
"sort": null,
"page": {
"size": 25
}
}Hugging Face Datasets2026 · Text
English and Yoruba Cassava SpeechEnglish and Yoruba Cassava Speech Dataset summary This dataset contains 200 short audio clips, totaling about 30.2 minutes: 50 questions and 50 answers in English, and the corresponding 50 questions and 50 answers in Yoruba. The topic is cassava cultivation, especially crop diseases, pests, planting, equipment, and processing. These are prepared prompts and answers recorded by two adults living in
Hugging Face Datasets2026 · Text · gated
Meddies ASR BenchMeddies ASR Bench This public dataset repository contains the validated 10 full audio chunks used for the MOSS + Gemini 3.8 Flash smoke benchmark. Contents data/full_10_chunks/audio/: ten full MP3 input chunks. data/full_10_chunks/moss/: the ten source MOSS-Diarize-Transcribe JSON records. data/full_10_chunks/manifest.json: chunk paths, offsets, durations, and embedded MOSS text. results/: the dow
Hugging Face Datasets2026 · Table · Parquet
Tutlait v1: Tamazight Speech with Arabic TranslationsTutlait v1: Tamazight Speech with Arabic Translations ~21 hours of spoken Tamazight (Amazigh / Berber) from 118 speakers, each clip paired with an Arabic (MSA) translation. It is built for speech-to-text translation (Tamazight audio → Arabic text) and for fine-tuning speech models such as Whisper on a low-resource language. Note: the text column is an Arabic translation of what was said, not a tra
Hugging Face Datasets2026 · Table · Parquet
Zimbabwe All SpeechZimbabwe All Speech Zimbabwe All Speech is an evolving, source-attributed collection of speech and paired text for Shona, Ndau, and Zimbabwean Ndebele. Teleagents brings together recordings from several hosted corpora with its own curation and data preparation. The collection is intended to make Zimbabwean local-language speech easier to find and use across research and model development. The rele
Hugging Face Datasets2026 · Table · Parquet
Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form)Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form) Spontaneous Lithuanian dialect speech from all four regions, with three transcripts per clip: written form (punctuation, capitalisation, numbers as digits), normalised, and the original phonetic transcription with stress marks. Dialect word forms are kept as spoken in every layer. text text_normalized text_phoneti
Hugging Face Datasets2026 · Table · Parquet
Multilingual LibriSpeech French (punctuated)Multilingual LibriSpeech French, punctuated and capitalized (train) A derivative of the French part of Multilingual LibriSpeech (MLS), the corpus of read audiobooks from LibriVox published by Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve and Ronan Collobert (Facebook AI Research). MLS distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upst
Hugging Face Datasets2026 · Text · gated
Hebrew Forced Alignment EvaluationHebrew Forced Alignment Evaluation Dataset Human-verified, word-level time-aligned Hebrew speech clips. To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was built. The system lets labelers fix the transcript and align each spoken word to the audio, down to 1ms precision (though annotators typically work at ~10ms granularity). The audio samples were gathered by
Hugging Face Datasets2026 · Table · Parquet
SANJAYKISHORE/TORGO-databaseThe TORGO Database: Acoustic and articulatory speech from speakers with dysarthria Dataset Summary This database only includes the short words and restricted sentence portion of the TORGO dataset. For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html. Transcripts have been normalized to remove punctua
Hugging Face Datasets2026 · Table · Parquet
IndicTelephony-BenchIndicTelephony-Bench v1.0 An open benchmark for speech recognition on Indian telephone speech: 25,393 human-curated utterances (30.2 hours) in nine languages, recorded over live phone lines and released exactly as the line delivered them, at 8 kHz. Most of it is code-mixed, much of it is short, and the business vocabulary a voice agent acts on is tagged in the references. At a glance Utterances 25
Hugging Face Datasets2026 · Text
Star Wars dialogue annotations (master cue table)Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with
Hugging Face Datasets2026 · Table · Parquet
ParlamentParla v3 (punctuated)ParlamentParla v3, punctuated and capitalized (train, short segments) A derivative of ParlamentParla v3, the speech corpus of Catalan parliamentary sessions published by the Language Technologies Unit of the Barcelona Supercomputing Center (BSC-LT) within the Aina project. ParlamentParla v3 distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text
Hugging Face Datasets2026 · dataset · gated
ST-NewsST-News 数据集卡片 / Dataset Card for ST-News 中文:ST-News(汕头新闻)是一个跨方言语音到字幕生成数据集,来源于汕头融媒集团新闻节目《今日视线》的播出档案。语料为潮汕话语音与普通话新闻字幕的配对数据——字幕遵循新闻字幕规范,并非逐字转写。本仓库为约 20 小时的公开发布子集,取自 2017 年播出数据。 English: ST-News (Shantou News) is a cross-dialect speech-to-subtitle generation dataset derived from broadcast archives of Jinri Shixian (《今日视线》), a news program of Shantou Media Convergence Group. It contains Teochew speech
Hugging Face Datasets2026 · Text
Voice Isolation Benchmark DatasetVoice Isolation Benchmark Dataset 265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing. 265 recordings · 47 speakers · 3 scenarios · 65 scripts Why this dataset exists Modern STT engines handle noise well. They still fail when a second person talks
Hugging Face Datasets2026 · dataset
GradrAI Viva Code-Switched Oral BenchmarkGradrAI Viva Code-Switched Oral Benchmark Consented, de-identified classroom-style oral answer clips used to benchmark GradrAI Viva for the Sahara CodeSwitch Africa challenge. Contents metadata.csv / metadata.jsonl: one row per clip. audio/: 16 kHz mono WAV files for Hugging Face dataset preview and ASR reuse. audio_original/: original submitted browser/Opus/WebM audio files. benchmark/: benchmark
Hugging Face Datasets2026 · Text · gated
Basis ConversationsDataset Card for Basis Conversations 1500 Listen first: sample conversations Overview Basis Conversations 1500 is a multi-party, multilingual, full duplex conversational speech dataset. Each conversation includes up to 4 simultaneous speakers, each with channel-separated, 48 kHz audio. The median conversation lasts 33 minutes and 2,645 unique speakers are represented. Multi-party: a conversation s
Hugging Face Datasets2026 · Table · Parquet
InteractSpeech English Text and TimelineInteractSpeech: English Text-and-Timeline Release This is the English text-and-timeline release associated with InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (Findings of EMNLP 2025). InteractSpeech is designed for real-time spoken-dialogue interaction, including interruptions, backchannels, pauses, gaps, overlaps, and turn transitions. The paper describes an appro
Hugging Face Datasets2026 · Table · Parquet · gated
DR P1 speech segmentsDR P1 speech segments Dataset Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors. Source The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565f
Hugging Face Datasets2026 · Image
ScreenASR-BenchScreenASR-Bench Data Item Value Split test Cases 2,002 Audio clips 2,002 Keyframes 2,469 Languages Chinese Structure Field Type Description caseid string Unique case identifier ref string Reference transcription target string Target text in TN form level string Difficulty level: L1, L2, or L3 audio audio Audio clip keyframes list[image] Keyframes associated with the case frame_captions list[string
Hugging Face Datasets2026 · Table · Parquet
atikuwu/karakalpak-speech-corpus📚 Karakalpak Speech Corpus (107 Hours) The Karakalpak Speech Corpus is the first comprehensive, open-access, community-crowdsourced speech recognition dataset for the Karakalpak language (kaa), a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan (Uzbekistan). Founded and led by Atabek Kadirbergenov alongside a student research team from the Muhammad al-Khwarizmi Speci
Hugging Face Datasets2026 · Table · Parquet · gated
KIRAAT — A Turkish Read-Speech CorpusKIRAAT — A Turkish Read-Speech Corpus A sentence-aligned read-speech corpus built from publicly available recordings on Turkish audiobook YouTube channels. The channel credits are in the table at the end of this card; every clip carries the channel it came from in the channel column. clips 1,840,404 duration 3,105.7 hours recommended subset 1,547,494 clips / 2,575.2 hours channels 27 speakers (clu
Hugging Face Datasets2026 · dataset · gated
ChaashiniChaashini (चाशनी) Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. Total: 1,391,986 clips · 2955.25 hours · 33 langu
Hugging Face Datasets2026 · Table · Parquet
Bambara stt ready for trainingOrigin : This is a ready-for-training clean bambara dataset good for STT models The RobotMali's bambara asr dataset (the most recent one) especially the bambara-asr-all subset has been cleaned and converted to 16 KHz freq. to avoid overfiting during training The the french translation column has be removed and too shorts (duration<0.9 sec) and too longs (duration > 30 secs) 's lines has been delet
Hugging Face Datasets2026 · Text
Open Yap 1K sampleOpen Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0,
Hugging Face Datasets2026 · Table · Parquet
AliAvd/persian-elderly-asrFinal gathered Persian elderly speech Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history. 80 paired recordings from four speaker folders. Reference transcripts were aligned
Hugging Face Datasets2026 · dataset · gated
Turkish Audiobook Speech Corpus (Raw)Turkish Audiobook Speech Corpus (Raw) Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir. İçerik Uzun-form Türkçe konuşma kayıtları (m4a / mp3) Kaynağa göre klasörlenmiş düz dizin yapısı Transkript, h