Constarium
← Search

Table · dataset · 2026

claimcheck-bench: Agent Success-Claim Verification

Listed in Hugging Face Datasets

claimcheck-bench A synthetic benchmark for checking whether an AI agent's success claim is supported by its tool results and environment state.

Description

An agent can say “done” after a failed write, an action on the wrong record, or an operation that never persisted. This dataset contains 300 labelled agent traces for evaluating detectors that distinguish successful completion from false success claims across booking, CRM, and coding tasks.

Created by Andrii Boiko as part of… See the full description on the dataset page: huggingface.co/datasets/aboiko/claimcheck-bench.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
text · text classification
From keywords
Computer Science & AI · Text
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsaboiko/claimcheck-bench6 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textmapping · Hugging Facevocabulary-mapper@1.0.0keywords['text']
concepts[task].hf_task:text-classificationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
licensesource · Hugging Faceconnector:huggingface@1.0.0/tags[license:*]
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0