Constarium
← Search

Table · dataset · 2022

Pangloss

Listed in Hugging Face Datasets

Dataset Card for [Needs More Information] Dataset

Description

Summary

Two audio corpora of minority languages of China (Japhug and Na), with transcriptions, proposed as reference data sets for experiments in Natural Language Processing. The data, collected and transcribed in the course of immersion fieldwork, amount to a total of about 1,900 minutes in Japhug and 200 minutes in Na. By making them available in an easily accessible and usable form, we hope to facilitate the development… See the full description on the dataset page: huggingface.co/datasets/Lacito/pangloss.

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Audio 75%
Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsLacito/pangloss12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:audiosource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:audioenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:automatic-speech-recognitionsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0