Constarium
← Search

Data · dataset · 2022

Broad Twitter Corpus

Listed in Hugging Face Datasets

This is the Broad Twitter corpus, a dataset of tweets collected over stratified times, places and social uses.

Description

The goal is to represent a broad range of activities, giving a dataset more representative of the language used in this hardest of social media formats to process. Further, the BTC is annotated for named entities.

For more details see https://aclanthology.org/C16-1111/

Links

Catalogue records · 1

Topics

Stated by source
token classification
Provenance · 1 source records, 9 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsstrombergnlp/broad_twitter_corpus11 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:token-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0