| Bibliography | Jin, Yucong: Do Machines Disagree like Humans in Ambiguous Visual Question Answering?. University of Stuttgart, Faculty of Computer Science, Electrical Engineering, and Information Technology, Bachelor Thesis No. 98 (2025). 42 pages, english.
|
| Abstract | Humans often give multiple valid but disagreeing answers to the same question about an image. To examine whether VQA models show similar behaviour, we study a transformerbased model: Vision-and-Language Transformer(ViLT), under two ambiguous cases: Synonym and Multi-reference. Using 3,990 question–image pairs from VQA-MHUG, which include human responses and gaze maps, we compare how often humans and machines disagree and how their attention pattern differs under the two ambiguous cases. Our results show that humans differ more both in answers and in their attention patterns, while the model remains more consistent. This takes a first step toward understanding how VQA models handle ambiguity and how their reasoning diverges from human behaviour.
|
| Department(s) | University of Stuttgart, Institute of Visualisation and Interactive Systems, Visualisation and Interactive Systems
|
| Superviser(s) | Bulling, Prof. Andreas; Kögel, Fabian; Hindennach, Susanne |
| Entry date | April 28, 2026 |
|---|