Imaging · dataset · 2026
VidScribe
Listed in Hugging Face Datasets
VidScribe VidScribe is a diagnostic benchmark for visual text in video generation.
Description
It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12).
Release status. This repository hosts the public half of the benchmark: 403 of 803 human-verified samples. The remaining samples will… See the full description on the dataset page: huggingface.co/datasets/Vicky0720/VidScribe.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/Vicky0720/VidScribe ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/Vicky0720/VidScribe ↗
metadata API · from Hugging Face
Topics
- Stated by source
- image · image to video · tabular · text · text to video · video · video to video
- From keywords
- Computer Science & AI · Optical character recognition
Provenance · 1 source records, 19 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | Vicky0720/VidScribe | 10 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[method].local:method:ocr | mapping · Hugging Face | vocabulary-mapper@1.0.0 | keywords['ocr'] |
| concepts[modality].hf_modality:image | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:tabular | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:video | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].local:modality:image | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[modality].local:modality:text | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[modality].local:modality:video | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[task].hf_task:image-to-video | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:text-to-video | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| concepts[task].hf_task:video-to-video | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license_text | source · Hugging Face | connector:huggingface@1.0.0 | |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |