Constarium
← Search

Table · dataset · 2025

FineWeb-10BT with WebOrganizer Labels for DoTA-RAG

Listed in Hugging Face Datasets

FineWeb-10BT with WebOrganizer Labels for DoTA-RAG Dataset · DoTA-RAG paper · Project page This dataset is a labeled version of the FineWeb-10BT web corpus used as the document collection in DoTA-RAG: Dynamic of Thought Aggregation RAG.

Description

It contains 14,868,862 English-language documents in one train split. Each document retains its FineWeb text and provenance fields and adds a predicted topic and document format from WebOrganizer.

The published Parquet files total 30.7 GB to… See the full description on the dataset page: huggingface.co/datasets/saksornr/fineweb-10bt-weborganizer-sigir-dota-rag.

Links

Documentation and papers

Catalogue records · 1

Topics

Stated by source
tabular · text · text retrieval
Inferred from text
Text 75%
Provenance · 1 source records, 12 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetssaksornr/fineweb-10bt-weborganizer-sigir-dota-rag12 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:tabularsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:text-retrievalsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0