INESC TEC Research Data Repository2026 · Table · CSV · unknown
Parallel Atomic and Hyper-Relational Reference Triple Annotations for European Portuguese Football Match ReportsThis dataset contains manually annotated reference triples derived from 50 professionally written Liga Portugal football match reports, created in support of the paper “A Writer–Editor–Critic Pipeline with Deterministic Fact Verification for Football Match Report Generation”, accepted at the 19th International Natural Language Generation Conference (INLG 2026). The annotations are provided in two
INESC TEC Research Data Repository2026 · Archive
Datasets and Models for Ontology-aligned Information Extraction in Portuguese Cultural HeritageThis repository contains datasets and fine-tuned models for ontology-aligned Named Entity Recognition (NER) and Relation Extraction (RE) in Portuguese cultural heritage archival documents. The datasets include annotations of entities and relations mapped to classes and properties from ArchOnto, an ontology designed for the Portuguese archives. The collection comprises both general-domain and domai
INESC TEC Research Data Repository2026 · Excel spreadsheet
Survey Dataset on LLM Adoption and Health Information Seeking: User Behaviour, Satisfaction, and Demographic InfluencesThis dataset was created to support the research study titled "A Complementary Tool for a Sensitive Domain: How Users Seek Health Information Through LLMs", which investigates how users integrate Large Language Models (LLMs) into their information-seeking process, with a particular focus on health-related contexts. Data were collected through an online survey, yielding a total of 405 responses, of
INESC TEC Research Data Repository2026 · Text
CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes - Version Archive**This record does not contain the dataset itself. It serves as a version archive for all published releases of CitiLink-Minutes.** This archive is the official reference for all releases of the CitiLink-Minutes dataset. It provides links to available dataset versions, each associated with its own DOI, documentation, annotation guidelines, and release notes. CitiLink-Minutes is a multilayer annota
INESC TEC Research Data Repository2026 · Text · unknown
CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes (v1.1.0)CitiLink-Minutes (v1.1.0) is a multilayer annotated dataset of 120 municipal meeting minutes from Portuguese city councils. Building on v1.0.0, this release extends the annotation coverage with: (1) personal information categorization, classifying anonymized spans into typed categories (PERSONAL-NAME, PERSONAL-ADDRESS, PERSONAL-ADMIN, PERSONAL-COMPANY, PERSONAL-POSITION, PERSONAL-PUBLIC); (2) hier
INESC TEC Research Data Repository2026 · Table · CSV
User Study Dataset for WikiQuality: A Browser Extension for Wikipedia Article Quality AssessmentThis dataset contains the questionnaires and anonymized responses from a user study evaluating WikiQuality, a Google Chrome extension for real-time Wikipedia article quality assessment. Participants compared pairs of Wikipedia movie articles in two conditions: without tool support and with the support of the extension. The dataset includes responses about article selection, decision accuracy, conf
INESC TEC Research Data Repository2026 · Text · unknown
CitiLink-Summ: A Dataset of Discussion Subjects Summaries in European Portuguese Municipal Meeting MinutesCitiLink-Summ is a domain-specific summarization dataset derived from the CitiLink-Minutes corpus, comprising 120 European Portuguese municipal meeting minutes from six municipalities. Each minute was manually segmented into discussion subjects and annotated with 2,880 high-quality, handwritten abstractive summaries (one per subject). The dataset supports research in segment-level automatic text s
INESC TEC Research Data Repository2026 · Image data
Taxonomy of SERP Elements Across Search EnginesThis dataset contains the dataset and supporting materials used in our research on Search Engine Results Pages (SERPs) structure, taxonomy, and element-level analysis across multiple search engines and languages. We used Python and Selenium Webdriver for web scraping the SERPs, ChatGPT-4o to translate the search queries. The dataset includes: - Multilingual search queries (English, Simplified Chin
INESC TEC Research Data Repository2026 · Text · unknown
LabadainLog-60: A Curated Tetun Query-Response Dataset for Conversational System Evaluation**1. Overview** LabadainLog-60 is a curated query dataset designed to evaluate and compare different conversational AI assistants for Tetun, derived from real user queries. It contains 60 queries selected from over 8,000 query log entries collected from two sources: - **Labadain Chat** (https://www.labadain.com): 2,101 entries collected from January 1–27, 2026 - **Labadain Old** (https://old.labad
INESC TEC Research Data Repository2025 · Archive
High-Resolution Clothing Segmentation Dataset for Deep LearningThis dataset contains images of garment items prepared for deep learning image segmentation tasks. The images were collected and organized to train, validate, and test segmentation models that can identify and isolate individual garments from the background. The data are structured into images and binary masks, but the masks can also be converted into annotation formats such as .txt or .json if re
INESC TEC Research Data Repository2025 · Text · unknown
CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes (v1.0.0)CitiLink-Minutes (v1.0.0) is a multilayer annotated dataset of 120 municipal meeting minutes from Portuguese city councils (31k annotations: 20,375 entities, 11,163 relations). It includes three annotation layers: **Metadata** (participants, location, date, meeting type), **subjects of discussion**, and **voting information** (subjects, positions, and outcomes). Data is provided in JSON format, on
INESC TEC Research Data Repository2025 · dataset · unknown
ClaimPT: A Dataset for Claim Detection and Fact-CheckingClaimPT is a Portuguese dataset of manually annotated claims in news articles. It contains 1,308 articles provided by the LUSA news agency, each annotated by two trained linguists following a dedicated scheme for claim detection. Annotations include claim spans, claimers, topics, stance, and temporal information. The resource supports research in claim detection, fact-checking, and misinformation
INESC TEC Research Data Repository2025 · Table · CSV
Interface Element Frequencies in Search Engine Results Pages (SERPs) Across Query Intents, Search Engines and LanguagesThis dataset contains the data produced for the dissertation "User Interface Variations in Search Engine Results Pages Across Types of Search Queries and Search Engines". The project was conducted by student Adelaide Miranda Santos at FEUP, University of Porto, as part of the Masters in Informatics and Computing Engineering. The primary objective of this work is to study interface variations in se
INESC TEC Research Data Repository2025 · Table · CSV
LusoClin: Dataset of synthetic clinical notes in European Portuguese generated using an open-source large language model, along with prompting and evaluation dataLusoClin is a publicly available dataset of fully synthetic clinical notes in European Portuguese, generated using an open-source large language model and carefully curated prompts. The dataset simulates realistic clinical narratives while ensuring that no real patient data is included, enabling privacy-preserving research on clinical text retrieval. The primary purpose of LusoClin is to support t
INESC TEC Research Data Repository2025 · Archive
Labadain-ZSRunS: Sparse and Zero-Shot Dense Retrieval Runs with LLM-Generated Summaries for Tetun Ad-Hoc Text Retrieval**1. Overview** Labadain-ZSRunS is a dataset consisting of run files produced by classical sparse and zero-shot dense retrieval models, resulted from the experiments on Tetun ad-hoc text retrieval. It also includes document summaries generated by a large language model (LLM) based on the full content of each document from the Labadain-Avaliadór collection. The dataset is intended to support resear
INESC TEC Research Data Repository2025 · Table · CSV
Labadain-Avaliadór : A Test Collection for Tetun Ad-hoc Text RetrievalThe Labadain-Avaliadór dataset is a test collection developed for the ad-hoc retrieval task. It comprises 59 topics, 33,550 documents, and 5,900 query-document relevance judgments (qrels), with an average of 36.76 relevant documents per query. The queries are sourced from real-world search activity, specifically from two channels: Google Search Console logs for Timor News and internal search logs
INESC TEC Research Data Repository2025 · Table · CSV
LabadainLog-17k+: Search Logs from Tetun-Speaking Users Across Chat, Web, and News Platforms**1. Overview** LabadainLog-17k+ is a dataset of interaction logs in Tetun, collected from three different platforms: - Labadain Chat (16,952 prompts): An LLM-powered conversational assistant tailored for Tetun speakers, accessible at www.labadain.com. - Labadain Search (400 queries): A monolingual search engine designed specifically for Tetun, available at www.labadain.tl. - Timor News (400 queri
INESC TEC Research Data Repository2025 · Text
Labadain-Stopwords: A Curated List of 160 Tetun StopwordsLabadain-Stopwords is a curated list of 160 Tetun stopwords, compiled from the Labadain-30k+ dataset and validated by native speakers. It is well-suited for various Tetun information retrieval and natural language processing tasks.The list is distributed in plain text format, with one word per line, enabling easy integration into various projects and applications.
INESC TEC Research Data Repository2025 · dataset
IILABS 3D: iilab Indoor LiDAR-based SLAM DatasetThe IILABS 3D dataset is a rigorously designed benchmark intended to advance research in 3D LiDAR-based Simultaneous Localization and Mapping (SLAM) algorithms within indoor environments. It provides a robust and diverse foundation for evaluating and enhancing SLAM techniques in complex indoor settings. The dataset was retrived in the Industry and Innovation Laboratory (iiLab) and comprises synchr
INESC TEC Research Data Repository2024 · Text
Raw atmospheric electric field and ancilllary data collected on-board Sagres ship (ongoing-updated yearly)This dataset comprises the raw atmospheric measurements collected on-board Sagres ship in the continuous monitoring campaign that followed the SAIL campaign, from May 2020 onwards. The SAIL campaign on board the iconic Portuguese tall ship NRP Sagres during its 2020 circumnavigation expedition from January to May 2020 aimed to improve the scientific understanding of the marine boundary layer throu
INESC TEC Research Data Repository2024 · dataset
Artificial Intelligence and Infodemic: Video Dataset for Fact-Checked Health Communication and Synthetic MediaVideos created using prototypes and APIs for participatory research. The videos were used as technological probes presented to various stakeholders. This dataset was created in the context of Fact-Checking Chatbot Initiative. The proliferation of disinformation poses a significant challenge to societies. Within the field of journalism, fact-checking emerges as a critical tool to combat this issue.
INESC TEC Research Data Repository2024 · Table · CSV
DECEiVeR (DatasEt aCting Emotions Valence aRousal)The DECEiVeR dataset (DatasetEt aCting Emotions Valence aRousal) is a meticulously curated collection of physiological recordings from 11 professional theatre actors. The DECEiVeR dataset facilitates the recognition of a specific set of five emotions: neutral, calm, tiredness, tension, and excitement. The dataset encompasses raw, low- and mid-level temporal and spectral data, enabling a comprehens
INESC TEC Research Data Repository2024 · Excel spreadsheet
RIS Based Hand Gesture Recognition DatasetThis dataset contains images for gesture recognition, divided into two main sets: dataset0608 and data_synthetic_variab. The data was collected using a wooden hand. **dataset0608** This dataset consists of two modes: ris_random and ris_optimized. The main difference between the two subfolders is the configuration of the RIS (random or optimized). This dataset consists of four subfolders: ris_rando
INESC TEC Research Data Repository2024 · Image data
Semantic representation of the Registos de Baptismos da Paróquia de Aldoar (Porto, Portugal)This dataset comprises mappings of archival records from the National Archives of Portugal to the RiC-O (Records in Contexts Ontology) framework, namely the baptism registries of the Parish of Aldoar (Porto, Portugal) (PT/ADPRT/PRQ/PPRT01/001). The original EAD XML file represents the raw archival data, which was subsequently processed through a data extraction script. This extraction parsed the o
INESC TEC Research Data Repository2024 · Table · CSV
Wikipedia and Simple Wikipedia Lead Section Pairs for Nine CategoriesThe dataset (categorized_dataset folder) contains 9 files in .csv format, each a collection of 10,000 lead section pairs sourced from Wikipedia (https://www.wikipedia.org/) and Simple Wikipedia (https://simple.wikipedia.org/) for a given category. Included categories are Culture, Education, Employment, Entertainment, Health, Leisure, Objects, Science and Time. This dataset was created to understan