| Bibliography | Feng, Xiwen: Survey and Evaluation of Explainable AI Methods for VQA Models that Learn from Disagreement. University of Stuttgart, Faculty of Computer Science, Electrical Engineering, and Information Technology, Bachelor Thesis No. 59 (2024). 36 pages, english.
|
| Abstract | Visual Question Answering (VQA) systems traditionally focus on identifying a single ”correct” answer for each image-question pair, neglecting the diversity of human interpretations. This study explores the phenomenon of disagreement in VQA, defined as variations in annotator-provided answers, and proposes a binary classification model to predict answer agreement. Using the VQA-MHUG dataset, visual features extracted with CLIP and textual features derived from DistilBERT are combined to train a Random Forest classifier. The model achieves a classification accuracy of 72.81%, demonstrating robust performance in identifying complete agreement but facing challenges in handling partial agreement cases. Additionally, a Local Interpretable Model-Agnostic Explanations (LIME)-based XAI framework is employed to interpret the model’s predictions, providing insights into the significance of image regions and question components. This work highlights the complexity of handling VQA disagreements and underscores the need for further advancements in feature integration and explanation methods to better capture human-centric variations in visual and textual understanding.
|
| Department(s) | University of Stuttgart, Institute of Visualisation and Interactive Systems, Visualisation and Interactive Systems
|
| Superviser(s) | Bulling, Prof. Andreas; Hindennach, Susanne |
| Entry date | February 21, 2025 |
|---|