Constarium
← Search

Data · dataset · 2025

USC-SFI MALACH Interviews and Transcripts English

Listed in UC Berkeley Library Dataverse

USC-SFI MALACH Interviews and Transcripts English – Speech Recognition Edition was developed by IBM as part of the MALACH (Multilingual Access to Large Spoken ArCHives) Project.

Description

This edition augments USC-SFI MALACH Interviews and Transcripts English (LDC2012S05) by modifying and updating a subset of the original corpus for use with the Kaldi toolkit in speech recognition work, and is easily portable for use by other speech recognition systems as well.

It contains approximately 168 hours of interviews from 682 Holocaust witnesses along with transcripts, a lexicon, Kaldi specific files, and other documentation. Inspired by his experience making Schindler’s List, Steven Spielberg established the Survivors of the Shoah Visual History Foundation in 1994 to gather video testimonies from survivors and other witnesses of the Holocaust. While most of those who gave testimony were Jewish survivors, the Foundation also interviewed homosexual survivors, Jehovah’s Witness survivors, liberators and liberation witnesses, political prisoners, rescuers and aid providers, Roma and Sinti (Gypsy) survivors, survivors of eugenics policies, and war crimes trials participants.

Read the rest (3 more)

The Foundation’s Visual History Archive holds nearly 55,000 video testimonies in 43 languages, representing 65 countries; it is the largest archive of its kind in the world. In 2006, the Foundation became part of the Dana and David Dornsife College of Letters, Arts and Sciences at the University of Southern California in Los Angeles and was renamed as the USC Shoah Foundation Institute for Visual History and Education.

The goal of the MALACH project was to develop methods for improved access to large multinational spoken archives; the focus was advancing the state of the art of automatic speech recognition and information retrieval. The characteristics of the USC-SFI collection -- unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching and emotional speech -- were considered well-suited for that task.

The work centered on five languages: English, Czech, Russian, Polish and Slovak. LDC has also released USC-SFI MALACH Interviews and Transcripts Czech (LDC2014S04).

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
Social Sciences
From keywords
Social Science
Inferred from text
Audio 65% · Video 75%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
UC Berkeley Library Dataversedoi:10.60503/D3/BIXEH310 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[field].dataverse_subject:social-sciencessource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/subjects
concepts[field].local:field:social-sciencemapping · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/subjects
concepts[modality].local:modality:audioenrichment · datasets lib berkeley edukeyword-concept-rules@1.0.0title+description (65%)
concepts[modality].local:modality:videoenrichment · datasets lib berkeley edukeyword-concept-rules@1.0.0title+description (75%)
created_datesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0
descriptionsource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/description
publication_datesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0
titlesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0/name
updated_datesource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0
version_labelsource · datasets lib berkeley educonnector:datasets_lib_berkeley_edu@1.0.0