Hugging Face Datasets2026 · Table · Parquet
Zimbabwe All SpeechZimbabwe All Speech Zimbabwe All Speech is an evolving, source-attributed collection of speech and paired text for Shona, Ndau, and Zimbabwean Ndebele. Teleagents brings together recordings from several hosted corpora with its own curation and data preparation. The collection is intended to make Zimbabwean local-language speech easier to find and use across research and model development. The rele
Hugging Face Datasets2026 · Text
Antalia evaluation suites and training-data manifestsAntalia evaluation suites and training-data manifests Companion data for Antalia 1 and Antalia 1 Foundation, an open Turkish text-to-speech release whose development is discontinued. No audio is included. The repository contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus filtering, the native-listening (CMOS) protocol with its raw single-listener results
Hugging Face Datasets2026 · Table · Parquet · gated
KIRAAT — A Turkish Read-Speech CorpusKIRAAT — A Turkish Read-Speech Corpus A sentence-aligned read-speech corpus built from publicly available recordings on Turkish audiobook YouTube channels. The channel credits are in the table at the end of this card; every clip carries the channel it came from in the channel column. clips 1,840,404 duration 3,105.7 hours recommended subset 1,547,494 clips / 2,575.2 hours channels 27 speakers (clu
Hugging Face Datasets2026 · dataset · gated
ChaashiniChaashini (चाशनी) Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. Total: 1,391,986 clips · 2955.25 hours · 33 langu
Hugging Face Datasets2026 · Text
Open Yap 1K sampleOpen Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial use Today we're releasing Open Yap 1K: 1,000 hours of dual-channel English conversation, capturing how people speak together naturally in real-world environments recorded in 48kHz. The dataset ships free for both commercial and research use. The sample on the Hugging Face Hub - 8.9 hours, 16 conversations, CC-BY-4.0,
Hugging Face Datasets2026 · Table · Parquet · gated
Datapoint Text-to-Speech Human Preferences (315K)Text-to-speech human preferences: 315K votes across 15 models This gated dataset contains the evaluation record behind Datapoint Audio Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech models in a complete round-robin over 300 English prompts. The prompt set covers eight practical voice-agent categories, and every generated sample is included as a typed audio record. The source ev
Hugging Face Datasets2026 · dataset · gated
Turkish Audiobook Speech Corpus (Raw)Turkish Audiobook Speech Corpus (Raw) Türkçe konuşma araştırmaları için derlenmiş, işlenmemiş uzun-form ses kayıtlarından oluşan bir koleksiyon. Kayıtlar çeşitli kaynaklardan bir araya getirilmiştir ve konuşmacı, kayıt ortamı, süre ve ses kalitesi bakımından geniş bir çeşitlilik gösterir. İçerik Uzun-form Türkçe konuşma kayıtları (m4a / mp3) Kaynağa göre klasörlenmiş düz dizin yapısı Transkript, h
Hugging Face Datasets2026 · Text
Quran Tajweed PhoneticsThe complete phonetic layer of the Quran in the riwaya of Hafs 'an 'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every phone carrying its tajweed attribution: madd class with its transmitted length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt, the seventeen sifat, and the rule that produced it. Built and maintained by Quran Lab, a waqf building open technology in
Hugging Face Datasets2026 · Table · Parquet
Ghana Speech — Audio with IPA TranscriptsGhana Speech — Audio with IPA Transcripts Speech with both transcript forms: the original orthography and the IPA phoneme sequence read off the audio by ASR. Each language is a subset, with real train/validation splits. from datasets import load_dataset ds = load_dataset("ghanaopendata/ghana-speech-ipa", "Akuapem_Twi_twi", split="train") ds[0]["audio"] # decoded waveform, 16 kHz ds[0]["text"] # or
Hugging Face Datasets2026 · Table · Parquet
YO-CPT-kkYO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities,
Hugging Face Datasets2026 · Table · Parquet · gated
Turkish TTS AudiobooksTurkish TTS Audiobooks Turkish read-speech corpus for text-to-speech training, built from Turkish audiobook and spoken-article recordings by an automatic pipeline: VAD segmentation → technical QC → acoustic event tagging → DNSMOS → speaker embedding/consistency → double-pass Whisper ASR → text policy → leakage-free splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards. The p
Hugging Face Datasets2026 · dataset
Pruna Skills doc examplesPruna Skills — doc examples Generated media and sidecars for PrunaAI/pruna-skills documentation. Each file under examples/ matches a skill demo in docs/EXAMPLES.md. PNG/MP3/MP4 outputs have a sibling .meta.json with the exact prompt, model, and inputs. Layout examples/ p-image-advanced.png p-image-advanced.meta.json quickstart-knight-still.png quickstart-knight-clip.mp4 chain-monarch-clip.mp4 … Re
Hugging Face Datasets2026 · Table · Parquet
Zenless VoiceZenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename
Hugging Face Datasets2026 · Table · Parquet
Artificial Voice Roleplay DatasetArtificial Voice Roleplay Dataset 67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character voice-direction captions with generated audio, across German, English, Spanish, and French (German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre, zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost
Hugging Face Datasets2026 · Text
Freya-TR-Eval (General Conversational Turkish TTS Benchmark)Freya-TR-Eval — A General-Purpose Turkish TTS Evaluation Benchmark Freya-TR-Eval is a compact, high-quality, reproducible benchmark of 495 natural conversational Turkish sentences for evaluating text-to-speech (TTS) systems on the everyday speech that conversational voice products actually synthesize. It is deliberately domain-neutral: it contains no bank-, call-center-, or product-specific conten
Hugging Face Datasets2026 · Table · Parquet
Common Voice Persian CleanPersian Common Voice Clean Dataset This dataset is a cleaned and prepared subset of the Persian (فارسی - fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12. The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences
Hugging Face Datasets2026 · Text
Vāgdhenu — Sanskrit Chant CorpusVāgdhenu — Sanskrit Chant Corpus A single-speaker Sanskrit chant (pārāyaṇa) recording corpus — classical ślokas chanted with tradition-faithful prosody and metrically-aware durations. Training data behind the Vāgdhenu Sanskrit Chant TTS. ~1,467 clips · ~5.3 hours · 24 kHz mono. One reciter (the author); classical śāstra chant, no Vedic svaras. Two configs (two cutting/normalization passes over the
Hugging Face Datasets2026 · Table · Parquet
LIEPA-3 Lithuanian Speech CorpusLIEPA-3 — Lithuanian Speech Corpus Didysis lietuvių kalbos garsynas (LIEPA-3) Dataset Summary LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours, ~7.5 million audio files) built for automatic speech recognition (ASR), text-to-speech (TTS) and linguistic research. It spans read, spontaneous, phonetically-annotated and dialectal speech recorded under a wide range of conditions (stud
Hugging Face Datasets2026 · Table · Parquet
espnet/yodas3YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 a
Hugging Face Datasets2026 · Table · Parquet
Burmese Synthetic Speech CorpusBurmese Synthetic Speech Corpus (DatarrX/burmese-synthetic-speech-corpus) Overview The Burmese Synthetic Speech Corpus is a high-fidelity, manually curated audio dataset specifically designed to advance Text-to-Speech (TTS) systems, speech recognition, and other audio-driven Machine Learning tasks for the Burmese (Myanmar) language. Created by DatarrX, this dataset bridges the gap in low-resource
Hugging Face Datasets2026 · Text
WorldSpeechWorldSpeech 🎉 WorldSpeech has been accepted to NeurIPS 2026! 🎉See the paper on arXiv. A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an
Hugging Face Datasets2026 · Table · Parquet · gated
burkimbia/sidbi-ziri-datasetSid Bi Ziri — Moore Language Speech Dataset Dataset Overview 56,992 audio segments of Moore language speech extracted from the Sid Bi Ziri TV show (SAVANE TV, Burkina Faso). Transcriptions were automatically generated by BIA-WHISPER (burkimbia/BIA-WHISPER-LARGE-SACHI_V2), a Whisper model fine-tuned on Moore. Note on transcription quality: This dataset is designed for TTS fine-tuning (SparkTTS). Si
Hugging Face Datasets2026 · Table · Parquet
Anilosan15/Synthetic_Turkish_TTS_DataSynthetic Turkish TTS Data This dataset was created by generating synthetic Turkish text across multiple speech scenarios. The text was produced in the following domains: finance_master, cs_master, parcel_delivery, ecommerce, telecom, isp_support, technical_support, subscription, insurance, health_appointments, public_services, education_registration, and daily_speech. These synthetic texts were t
Hugging Face Datasets2026 · dataset
Tarifit TTS CorporaTarifit TTS Corpora First phonetically balanced corpora for Tarifit (Riffian Berber). Designed for Text-To-Speech (TTS) training. Files File Description Sentences tarifit_pbc_corpus.csv Phonetically Balanced Corpus with IPA 57 metadata.csv Coqui TTS training format 57 tarifit_customer_service_tts.csv Native-validated customer service corpus 24 Language Tarifit — Riffian Berber (ISO 639-3: rif) Spo
Hugging Face Datasets2026 · Text · gated
DMC-ykfx33/nsfw_tts_datasetA high-quality audio dataset designed for training and fine-tuning NSFW TTS models, including 30 characters, over 1000 hours of audio, and rich emotion/sound annotations. Sample format: WAV (audio) + TXT (annotations), including emotion_label, sound_label and text. Annotations: 6000+ emotion labels (intimate, breathy, teasing, etc.) and 760+ sound labels (moan, sigh, laugh, etc.) in the full versi