Data · dataset · 2025
MASBA: A Large-Scale Dataset for Multi-Level Abstractive Summarization of Bangla Articles
Listed in Teesside University Research Data Repository
Our research hypothesis is to evaluate the effectiveness of different Bangla text summarization methods compared to the original text ('main').
Description
The data shows that: - The average length of the main text is 2482.72 characters. - The average length of the summaries are: - sum1: 293.75 characters, - sum2: 506.10 characters, - sum3: 688.50 characters. The compression ratio of each summary method (summary length divided by main length) reveals that: - sum1's mean compression ratio is 0.14, - sum2's mean compression ratio is 0.24, and - sum3's mean compression ratio is 0.33.
Notable findings: - sum1 appears to be the shortest summary on average, with a higher degree of compression. - sum2 produces summaries of medium length, while sum3 tends to generate the longest summaries. Data Gathering and Interpretation: The data can be interpreted to assess which method produces the most concise, yet meaningful, summaries. Researchers can use these findings to evaluate the trade-offs between summary length and completeness of information conveyed.
Links
Where it is published
- DOI doi.org/10.17632/rxhj7g6y2k.3 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
- From keywords
- Computer Science & AI · Earth & Environmental Science · Engineering · Humanities · Life Sciences · Natural language processing · Social Science
- Inferred from text
- Text 75%
Provenance · 1 source records, 13 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/rxhj7g6y2k.3 | 9 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].anzsrc:field:460208 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Natural Language Processing'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/rights |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |