Large vision-language models can describe photographs of dogs, read text on street signs, and answer questions about charts. But ask one to interpret a painting, explain why a character looks anxious, or identify the cultural tradition behind a temple mural, and the picture changes. These models are entering classrooms and tutoring systems where the visual content is not a snapshot of everyday life but artwork, illustration, and culturally situated imagery. The gap between what current benchmarks measure and what educational deployments require is large, and largely uncharted.
A team from AI Singapore, the National University of Singapore, and A*STAR built MUSE to map that gap. The benchmark evaluates 30 vision-language models across 12 tasks designed around artistic imagery used in language education. The tasks span visual perception, semantic understanding, affective interpretation, compositional reasoning, and cultural understanding, with images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions.
Why existing benchmarks miss the point for education
Current VLM benchmarks, including MMBench, SEED-Bench, MMMU, and MathVista, evaluate models on natural photographs, diagrams, charts, and examination materials. They test whether a model can recognize objects, read text, or solve domain-specific reasoning problems. These are useful capabilities, but they do not capture what a model needs to do when a student shows it a painting and asks "Why does this person look sad?" or "What culture does this image represent?"
Artistic imagery introduces complications that natural photographs do not. Paintings use stylized forms, non-photorealistic colors, and implicit narratives. A character's emotion might be conveyed through brushstroke direction, color palette, and spatial composition rather than through a facial expression that a classifier trained on photographs can read. Cultural elements, such as clothing, architectural details, or symbolic motifs, require knowledge that extends beyond object recognition into domain-specific cultural literacy.
Existing art-focused benchmarks like ArtEmis, ArtELingo, VQArt-Bench, and AICA-Bench each address a piece of this puzzle, but none jointly evaluate the full range of capabilities an educational VLM needs. They focus on individual domains, affective understanding in isolation, or cultural QA without compositional reasoning. MUSE is designed to test all five dimensions together, using the same set of images, so that a model's performance on one capability can be compared directly against its performance on others.
Decoupling annotation from question generation
The benchmark's construction is its most methodologically interesting feature. MUSE separates image annotation from question generation. Each artwork is annotated once with a reusable structured representation of its visual and semantic content: bounding boxes for characters, objects, and activities; labels for emotions, scenes, and cultural elements; descriptions of activities; counts of objects; and causal explanations for detected emotions. These annotations are the single source of truth for all 12 tasks.
Task-specific questions are then instantiated through predefined generation rules that consume the shared annotation. This means adding a new task does not require re-annotating images. It requires writing a new generation rule that queries the existing annotations in a different way. The design also makes question difficulty controllable, because the generation rules can select which aspects of the annotation to query and how many distractors to include.
This is a practical response to a real problem in benchmark construction. Conventional pipelines annotate images independently for each task, producing task-specific question-answer pairs that cannot be reused. Scaling to 12 tasks with hundreds of questions each would require enormous annotation effort if done the traditional way. By annotating once and generating many times, MUSE achieves 2,400 questions over 1,174 images with a fraction of the per-task annotation cost.
The annotation process itself involved 127 annotators who received a briefing on the study motivation, task definitions, guidelines, and representative examples. Each sample was independently annotated by one annotator, reviewed by two others, and finalized only after consensus, with disagreements resolved through the established guidelines.
Twelve tasks across five capability dimensions
The 12 tasks are organized into five dimensions that reflect the capabilities an educational VLM needs.
Visual Perception covers object classification (given a bounding box, classify the character) and object counting (numerically predict object counts, testing recognition under occlusion and size variation).
Semantic Understanding includes scene classification (integrate global visual and semantic information to classify the overall scene), activity localization (select the bounding box corresponding to a described activity from among distractors), and activity description (identify the correct description of what is happening in a bounding box, with 10 different distractor generation methods including cross-image negatives and concatenated activity descriptions).
Affective Interpretation forms a four-turn sequence: emotion detection (classify a character's emotion using categories from Plutchik's emotion wheel), visual clue identification (open-ended description of the visual evidence supporting the emotion prediction), and emotion cause inference (open-ended explanation of the emotion's cause). The first two tasks in this sequence produce multiple-choice questions, while the latter two are open-ended and evaluated by semantic similarity between model output and human references.
Compositional Reasoning includes relative position (predict three-dimensional spatial relations from the characters' viewpoints: left/right, front/back, above/below), remote interaction (reason about non-contact interactions between entities, requiring both identification of the interacting entity and supporting visual evidence), and jigsaw puzzle (complete puzzles by aligning patches through continuity in shape, color, and texture, using five different segmentation grids and various piece shapes).
Cultural Understanding includes cultural identification (identify cultural elements within a bounding box, with 15 clustered categories derived from OpenAI text embeddings).
Most tasks use textual or visual multiple-choice questions. Object count requires numerical prediction. Visual clue identification and emotion cause inference use open-ended responses evaluated by semantic similarity (cosine similarity between text-embedding-3-large embeddings).
What the evaluation reveals about 30 models
The evaluation covers 30 models spanning open-source and proprietary systems: GPT-5.6-Sol, GPT-4o, Qwen3-VL-32b, InternVL3-38b, Qwen2.5-VL-72b, Qwen3-VL-8b, InternVL3-14b, and many others across multiple families including Gemma 3, MiniCPM-V, DeepSeek-VL2, CogVLM2, GLM-4V, LLaVA-NeXT, and Yi-VL.
GPT-5.6-Sol achieves the strongest overall results, surpassing GPT-4o on 10 of 12 tasks. But the picture is more nuanced than a single leader. GPT-4o remains superior on activity description and relative position. Qwen3-VL-32b leads on activity localization and scene classification. InternVL3-14b performs best on relative position. GLM-4V-9b leads on jigsaw puzzle. No model dominates across all tasks.
The performance divide between recognition and integrative reasoning is stark. Scene classification is comparatively mature, with 23 of 30 models exceeding 75.0 and a median score of 81.0. Emotion detection, relative position, remote interaction, and jigsaw puzzle exhibit substantially lower medians. The best models score 39.5 on emotion detection, 10.0 on relative position, 86.5 on remote interaction (but most models score far lower), and 42.0 on jigsaw puzzle. Models are reliable at recognizing what is visibly present but struggle to ground predictions in visual evidence, explain affective causes, or reason over perspective-dependent relations.
Scaling is non-monotonic. InternVL3-38b outperforms InternVL3-14b on most tasks but performs worse on activity description and relative position. This suggests that simply making a model larger does not uniformly improve all capabilities. Architecture and training choices remain important determinants of capability-specific performance.
Affective understanding is not a single skill
The affective computing analysis reveals that emotion detection, visual clue identification, and emotion cause inference do not form a unified capability. A model that classifies emotions well does not necessarily ground its predictions in visual evidence or explain the causes correctly. Performance rankings shift across these three tasks, indicating that affective understanding in artistic imagery requires a chain of capabilities that current models have not learned to integrate.
The error analysis shows a cascading failure pattern. When a model misidentifies the target character (for example, shifting attention from the intended character to a salient nearby figure), it then predicts the wrong emotion for the wrong character, and constructs a coherent explanation for that incorrect prediction using visual elements that actually belong to the correct character. The failure at the grounding stage propagates through the entire reasoning chain. GPT-5.6-Sol correctly identifies the target in one example but misreads a stylized expression as surprise, indicating an affect-interpretation error rather than a grounding failure. Other models often attend to a different character entirely, predict joy, and justify the prediction using butterflies, birds, or nearby interactions that are visually present but irrelevant to the actual target.
This is particularly problematic for educational applications. An AI tutor that misidentifies which character is expressing an emotion, then constructs a plausible explanation for the wrong character, would give students incorrect feedback with high confidence. The model's output reads as authoritative and well-reasoned, but it is fundamentally misaligned with the visual evidence.
Spatial reasoning remains a persistent bottleneck
The relative position task exposes a forced-relation bias in spatial reasoning. When the ground truth specifies no definite lateral or vertical relation between two characters, 90.0% of models predict one anyway. Depth reasoning is also unreliable, with only 43.3% of models correctly identifying which character is in front. No model resolves all three spatial dimensions correctly in the tested examples.
Current VLMs oscillate between two failure modes: asserting definite relations under ambiguous evidence, or predicting None across all dimensions and missing valid depth cues. This reveals weak viewpoint-aware spatial reasoning and poor calibration of spatial uncertainty. For educational use, this matters when a model needs to describe the spatial relationships in a scene, such as explaining where characters stand relative to each other or what spatial composition the artist used.
Correlation analysis shows MUSE measures distinct capabilities
Spearman rank correlations across 30 models reveal that the 12 MUSE tasks form related but non-redundant capability groups. Object count, emotion detection, and scene classification are strongly correlated. Visual clue identification closely tracks emotion cause inference, linking visual evidence grounding with affective reasoning. Activity localization, activity description, and remote interaction form a moderately correlated group centered on entity-activity and cross-region reasoning. Jigsaw puzzle correlates weakly with most other tasks, indicating a distinct compositional capability.
When compared against six existing benchmarks, MUSE shows partial but incomplete overlap. MMBench and AI2D correlate strongly with several MUSE tasks, indicating shared perceptual and semantic capabilities. BLINK exhibits inconsistent correlations. Activity description and relative position assess capabilities underrepresented in existing benchmarks. AICA-Bench emotion reasoning aligns moderately with MUSE's affective tasks, suggesting that emotion reasoning on conventional visual content only partially transfers to stylized expressions and implicit narratives in artworks.
The practical implication is that a model's score on MMBench does not predict its ability to interpret artistic imagery in educational settings. Developers and educators need to evaluate models on the actual content they will encounter, not assume that performance on natural photographs transfers.
Limitations and open questions
MUSE focuses on artistic imagery for language education. The 12 tasks are designed around the specific capabilities needed for image-based learning, not for general art history, art criticism, or creative applications. The images are curated to center Singaporean and Southeast Asian contexts alongside Western art, which is deliberate for cultural diversity but limits generalizability to other cultural traditions.
The annotation-first, task-generative design is scalable but depends on the quality and completeness of the initial annotations. If an annotator misses a visual clue or mislabels an emotion, the error propagates to all tasks that query that annotation. The quality control process, with independent annotation and consensus resolution, mitigates this but does not eliminate it.
The evaluation uses temperature 0 and evaluates each model once, which the authors justify by noting output stability. But temperature 0 may not reflect how models are deployed in practice, where sampling introduces variability that could affect performance on open-ended tasks like visual clue identification and emotion cause inference.
The semantic similarity evaluation for open-ended tasks, using cosine similarity between text-embedding-3-large embeddings, is a reasonable proxy but not a perfect measure of answer quality. A model that produces a semantically similar but factually incorrect explanation would receive a high score.
What this means for building educational VLMs
MUSE provides a standardized benchmark for measuring progress on the capabilities that educational VLMs actually need. The results show that current models are good at recognizing what is present in an image but poor at explaining why, grounding their explanations in visual evidence, reasoning about spatial relationships, and identifying cultural content. Scaling alone does not solve these problems. What is needed is region-aware grounding, integrated reasoning from perception to evidence to causes, and viewpoint-aware modeling of spatial uncertainty.
For developers, the benchmark offers a concrete evaluation target. The 12 tasks and five capability dimensions define what "good" looks like for educational VLMs. The annotation-first construction framework means the benchmark can be extended with new tasks and new images without re-annotating everything from scratch. The dataset and code are publicly available on Hugging Face.
For educators considering AI-assisted tutoring, the results are a reality check. A model that performs well on general VLM benchmarks may still fail at the interpretive and reasoning tasks that educational settings demand. Evaluating models on MUSE before deployment would catch these gaps before they affect students.
Read the paper on arXiv