Constarium
← Search

Text · dataset · 2022

MuchMore

Listed in Hugging Face Datasets

The corpus used in the MuchMore project is a parallel corpus of English-German scientific medical abstracts obtained from the Springer Link web site.

Description

The corpus consists approximately of 1 million tokens for each language. Abstracts are from 41 medical journals, each of which constitutes a relatively homogeneous medical sub-domain (e.g. Neurology, Radiology, etc.).

The corpus of downloaded HTML documents is normalized in various ways, in order to produce a clean, plain text version, consisting of a title, abstract and keywords. Additionally, the corpus was aligned on the sentence level. Automatic (!) annotation includes: Part-of-Speech; Morphology (inflection and decomposition); Chunks; Semantic Classes (UMLS: Unified Medical Language System, MeSH: Medical Subject Headings, EuroWordNet); Semantic Relations from UMLS.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text
Inferred from text
Text 75%
Provenance · 1 source records, 10 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsbigbio/muchmore12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0