Constarium
← Search

Table · dataset · 2026

VLADBench-reeval

Listed in Hugging Face Datasets

VLADBench-reeval We re-evaluated the VLADBench benchmark (Li et al., 2025, arXiv:2503.21505) against current SOTA VLMs under the original scoring criteria and prompts.

Description

See Eventual-Inc/VLADBench for the code, the companion site for interactive results, and the article for a summary of the findings. Cost vs Score Up and to the left is better.

The dashed line is the cost-performance frontier which is defined by measuring the best TOTAL score available against its… See the full description on the dataset page: huggingface.co/datasets/Eventual-Inc/VLADBench-reeval.

Links

Documentation and papers

Catalogue records · 1

Topics

Stated by source
tabular · text · visual question answering
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsEventual-Inc/VLADBench-reeval9 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:tabularsource · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].hf_modality:textsource · Hugging Faceconnector:huggingface@1.0.0
concepts[task].hf_task:visual-question-answeringsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0