Monica Riedler
Many AI benchmarks assume each example has a single correct label. In practice, however, people often disagree even on the same input. This phenomenon, known as Human Label Variation (Plank, 2022), suggests that annotator disagreement is not simply noise but can reflect meaningful differences in how people interpret data (Aroyo and Welty, 2015; Pavlick and Kwiatkowski, 2019; Uma et al., 2022). In language tasks, such variation can arise from ambiguity, context, background knowledge, or differences in perspective. In vision tasks, it can reflect attention to different perceptual cues or image features. Vision-language tasks bring these sources of variation together, requiring people not only to interpret images and text, but also to determine how the two relate to each other. This makes multimodal settings a particularly relevant context for studying human disagreement, as meaning is often shaped by the interaction of linguistic, visual, and contextual cues. However, multimodal foundation models are still commonly evaluated against single ground-truth labels, which can hide plausible variation in human judgments and encourage models to favor one dominant interpretation. As a result, current evaluation practices make it difficult to assess whether models are aligned with the range of human perspectives, or primarily optimized toward simplified and potentially biased labels. Although Human Label Variation has been studied in NLP and, to some extent, in computer vision, its role in vision-language tasks remains underexplored. This project addresses this gap by investigating how human disagreement arises in multimodal settings, using both labels and annotators' natural-language explanations to examine the cues and reasoning strategies behind different judgments. These data will be used to analyze where multimodal models reflect plausible human variation, where they fail to capture diverse perspectives, and where they make brittle, biased, or poorly calibrated predictions. Building on these insights, the project will develop HLV-aware resources and methods for evaluating and adapting multimodal systems. Its broader aim is to support models that are better calibrated, more robust, and more inclusive by aligning them not only with majority judgments, but also with the diversity of human interpretation.