Constarium
← Search

Data · dataset · 2026

exploitbench/v8

Listed in Hugging Face Datasets

ExploitBench V8 — v8-codex-ace-83a40e1-ptf81548b Per-cell exploitation results from the V8 JavaScript engine benchmark, with full transcripts, tool-call logs, and capability grading.

Description

This dataset is the academic record for ExploitBench: succeeded runs and model-failed runs both ship, including cells where the model gamed the grader (see audit.json). Envs in this revision 41 environments.

Full list — one per env_id, sorted: v8-crbug-1509576 v8-crbug-339064932… See the full description on the dataset page: huggingface.co/datasets/exploitbench/v8.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
reinforcement learning
Provenance · 1 source records, 8 field assertions
SourceKeyLast seenRaw
Hugging Face Datasetsexploitbench/v88 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:reinforcement-learningsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0