figshare + Loughborough Research Repository + GRANTS Data + UP Research Data Repository2026 · Astronomical catalogue
Consonant articulating errors in Khmer preschoolers (Linna & Zhou, 2026)<p dir="ltr"><b>Purpose: </b>Consonant errors (mainly phonetic and phonological errors) as a phenomenon in speech development are frequently observed in preschool children’s daily speech communication across languages. Since the Khmer language is a low-source language, consonant articulating error patterns of Khmer-speaking children have not been explored to date, leaving open the question whether
ZivaHub + Deakin Research Online + DMU Figshare + HKU DataHub + Swinburne Figshare + DaYta Ya Rona + SUNScholarData + figshare + Loughborough Research Repository + GRANTS Data + UP Research Data Repository2026 · Astronomical catalogue
Chesapeake Bay Island Bird Audio, Transect Data, and Code for Nestedness/Modularity Analysis<p dir="ltr">Metacommunity structure has often been characterized in terms of the degree of nestedness (i.e., when each community’s composition is a subset of more speciose communities) and the degree of modularity (i.e., when species form discrete groups of co-occuring taxa that do not overlap). For a conservation biologist interested in preserving all species, a patch with the most species-rich
Hugging Face Datasets2026 · Text
English and Yoruba Cassava SpeechEnglish and Yoruba Cassava Speech Dataset summary This dataset contains 200 short audio clips, totaling about 30.2 minutes: 50 questions and 50 answers in English, and the corresponding 50 questions and 50 answers in Yoruba. The topic is cassava cultivation, especially crop diseases, pests, planting, equipment, and processing. These are prepared prompts and answers recorded by two adults living in
Hugging Face Datasets2026 · Text
Hamozwa/RepeatRealDataset Card for RepeatReal Real-world portion of the audio datasets referenced in Class-Agnostic Audio Repetition Counting. Code: https://github.com/Hamozwa/audio-counting Dataset Summary Unlike the synthetic RepeatSynth data, these are real, variable-length recordings across three domains, used as held-out test sets for models trained only on synthetic data. Config Domain Source clocks Mechanica
Hugging Face Datasets2026 · Image
MiniMax-H3 video benchmark mediaMiniMax-H3 video benchmark media Reference images, videos and audio for the MiniMax-H3 video benchmark, hosted so an inference server can fetch them by URL: https://huggingface.co/datasets/zhenghaoniTT/minimax-h3-bench-assets/resolve/main/<file> Attribution Videos, audio and the Tears of Steel stills are derived (trimmed, cropped, re-encoded) from Tears of Steel, (c) copyright Blender Foundation |
Hugging Face Datasets2026 · Text · gated
Meddies ASR BenchMeddies ASR Bench This public dataset repository contains the validated 10 full audio chunks used for the MOSS + Gemini 3.8 Flash smoke benchmark. Contents data/full_10_chunks/audio/: ten full MP3 input chunks. data/full_10_chunks/moss/: the ten source MOSS-Diarize-Transcribe JSON records. data/full_10_chunks/manifest.json: chunk paths, offsets, durations, and embedded MOSS text. results/: the dow
Hugging Face Datasets2026 · Table · Parquet
Tutlait v1: Tamazight Speech with Arabic TranslationsTutlait v1: Tamazight Speech with Arabic Translations ~21 hours of spoken Tamazight (Amazigh / Berber) from 118 speakers, each clip paired with an Arabic (MSA) translation. It is built for speech-to-text translation (Tamazight audio → Arabic text) and for fine-tuning speech models such as Whisper on a low-resource language. Note: the text column is an Arabic translation of what was said, not a tra
Hugging Face Datasets2026 · Table · Parquet
Zimbabwe All SpeechZimbabwe All Speech Zimbabwe All Speech is an evolving, source-attributed collection of speech and paired text for Shona, Ndau, and Zimbabwean Ndebele. Teleagents brings together recordings from several hosted corpora with its own curation and data preparation. The collection is intended to make Zimbabwean local-language speech easier to find and use across research and model development. The rele
Hugging Face Datasets2026 · Text
AIMS-RAIL/RAILRAIL Audio Benchmark This folder is generated for direct Hugging Face dataset upload. Each row uses relative audio paths rooted at this repository folder. NeurIPS / Croissant Metadata metadata.json is a Croissant-style metadata file with core fields and minimal RAI fields. metadata.json includes the Hugging Face dataset URL, CC-BY-4.0 license URL, checksums, and RAI fields. build_summary.json cont
Hugging Face Datasets2026 · Table · Parquet
Hamozwa/RepeatSynthDataset Card for RepeatSynth Synthetic portion of the audio datasets referenced in Class-Agnostic Audio Repetition Counting. Code: https://github.com/Hamozwa/audio-counting Dataset Summary RS, RSN, and RVN are synthetic datasets containing varying levels of background noise. Every sample is uniformly 10 seconds long and contains between 0 and 8 repetition events. Config Description rs RepeatSound
Hugging Face Datasets2026 · dataset · gated
mten223/q7x9v2k4m8p1z6r3n5c0w7t9b2h4f8d1s6j3y0e5u9i2o7a4g8l1Hugging Face Datasets2026 · Text
dacvae-tts Turkish 115M: trc-w640 (tr-combined + tr-dataset-12)dacvae-tts Turkish 115M: trc-w640 (tr-combined + tr-dataset-12) Checkpoints of the training run trc-w640, each with 5 outputs: Turkish zero-shot voice-cloning TTS (dacvae-tts, branch turkish-tts; flow-matching DiT on frozen Meta DACVAE latents, 48 kHz). Training in progress (update 61200 of 100000): every 10000 updates a new checkpoint and its 5 outputs are added here. 5 outputs per checkpoint The
Hugging Face Datasets2026 · Table · Parquet
Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form)Lithuanian Dialect Speech 100 h: punctuated, cased, numbers as digits (written form) Spontaneous Lithuanian dialect speech from all four regions, with three transcripts per clip: written form (punctuation, capitalisation, numbers as digits), normalised, and the original phonetic transcription with stress marks. Dialect word forms are kept as spoken in every layer. text text_normalized text_phoneti
Hugging Face Datasets2026 · Table · Parquet
Multilingual LibriSpeech French (punctuated)Multilingual LibriSpeech French, punctuated and capitalized (train) A derivative of the French part of Multilingual LibriSpeech (MLS), the corpus of read audiobooks from LibriVox published by Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve and Ronan Collobert (Facebook AI Research). MLS distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upst
Hugging Face Datasets2026 · Text · gated
Hebrew Forced Alignment EvaluationHebrew Forced Alignment Evaluation Dataset Human-verified, word-level time-aligned Hebrew speech clips. To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was built. The system lets labelers fix the transcript and align each spoken word to the audio, down to 1ms precision (though annotators typically work at ~10ms granularity). The audio samples were gathered by
Hugging Face Datasets2026 · Table · Parquet
SANJAYKISHORE/TORGO-databaseThe TORGO Database: Acoustic and articulatory speech from speakers with dysarthria Dataset Summary This database only includes the short words and restricted sentence portion of the TORGO dataset. For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html. Transcripts have been normalized to remove punctua
Hugging Face Datasets2026 · dataset
DACVAE-TTS Turkish run C (width 512, clean data)DACVAE-TTS Turkish run C (width 512, clean data) Generated audio of every evaluated checkpoint of the training run tr-w512-clean (Turkish zero-shot voice-cloning TTS, dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments). Run C: configs/nano_tr_w512.yaml (66.5M parameters: wi
Hugging Face Datasets2026 · Table · Parquet
IndicTelephony-BenchIndicTelephony-Bench v1.0 An open benchmark for speech recognition on Indian telephone speech: 25,393 human-curated utterances (30.2 hours) in nine languages, recorded over live phone lines and released exactly as the line delivered them, at 8 kHz. Most of it is code-mixed, much of it is short, and the business vocabulary a voice agent acts on is tagged in the references. At a glance Utterances 25
Hugging Face Datasets2026 · Text
ODU-BenchODU-Bench: Omni Demand Understanding Introduction ODU-Bench is a benchmark for contextual user-intent inference in audio and audio-visual interaction. Omni Demand Understanding (ODU) asks a model to detect whether a valid user demand is present and infer the user's intent from multimodal and conversational context. A demand is the outcome a user wants an assistant to achieve, including response-re
Hugging Face Datasets2026 · Table · Parquet
DeepASMR-NSpeechDeepASMR-NSpeech A Fine-Grained Benchmark for Non-Speech ASMR Generation 73,829 ten-second clips · 202.6 hours · 37 fine-grained actions 🎧 Interactive Demo · ⌘ Code · 📄 Paper: coming soon Dataset summary DeepASMR-NSpeech is a non-speech ASMR audio dataset with structured Subject-Verb-Object annotations and the SVO-AQA audio question-answering benchmark. All clips are 10 seconds long and cover 37 f
Hugging Face Datasets2026 · Table · Parquet
ParlamentParla v3 (punctuated)ParlamentParla v3, punctuated and capitalized (train, short segments) A derivative of ParlamentParla v3, the speech corpus of Catalan parliamentary sessions published by the Language Technologies Unit of the Barcelona Supercomputing Center (BSC-LT) within the Aina project. ParlamentParla v3 distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text
Hugging Face Datasets2026 · dataset · gated
otoSpeech-full-duplex-task-oriented-v2: Full-Duplex Task-Oriented ConversationsDataset Viewer https://cc-task-oriented-preview.vercel.app/ otoSpeech-full-duplex-task-oriented-v2 Contact This sample dataset is provided for research purposes. We maintain larger and more diverse datasets. For collaborations, inquiries, custom data collection, or joint research, contact consome@oto.earth. Dataset Summary otoSpeech-full-duplex-task-oriented-v2 contains 17.376131 hours of English,
Hugging Face Datasets2026 · Text
Voice Isolation Benchmark - Processed AudiosVoice Isolation Benchmark – Processed Audios Audio samples processed by four Krisp Voice Isolation models. Each scenario folder contains subfolders for every model, with one processed .wav file per original sample. Voice Isolation Models VI 2.5 Default (vi_2_5_default) The main Voice Isolation model. Strongest at removing noise and other speakers. Goes fully silent when only a bystander is talking
Hugging Face Datasets2026 · Image · gated
gulucaptain/MiniMax-H3-ReasonMiniMax-H3-Reason: Evaluation Data This directory contains the local preparation scaffold for the benchmark release. Actual evaluation inputs and prompts have not yet been populated here. Files under templates/ are organization templates, not benchmark samples. Input modalities Every instance includes a text prompt. Scenario Media inputs MSR One image ADR One image and one audio clip VDR One video
Hugging Face Datasets2026 · dataset · gated
ST-NewsST-News 数据集卡片 / Dataset Card for ST-News 中文:ST-News(汕头新闻)是一个跨方言语音到字幕生成数据集,来源于汕头融媒集团新闻节目《今日视线》的播出档案。语料为潮汕话语音与普通话新闻字幕的配对数据——字幕遵循新闻字幕规范,并非逐字转写。本仓库为约 20 小时的公开发布子集,取自 2017 年播出数据。 English: ST-News (Shantou News) is a cross-dialect speech-to-subtitle generation dataset derived from broadcast archives of Jinri Shixian (《今日视线》), a news program of Shantou Media Convergence Group. It contains Teochew speech