Constarium
← Search

Data · dataset · 2026

Chaashini

Listed in Hugging Face Datasets

Description

Chaashini (चाशनी) Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. Total: 1,391,986 clips · 2955.25 hours · 33 languages Format: mono 24 kHz FLAC (audio column) with a verbatim transcript and rich… See the full description on the dataset page: huggingface.co/datasets/kapturecx/Chaashini.

Links

Get the data

Catalogue records · 1

Topics

Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetskapturecx/Chaashini7 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audiomapping · Hugging Facevocabulary-mapper@1.0.0keywords['speech']
concepts[task].hf_task:audio-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:automatic-speech-recognitionsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-to-speechsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0