Master Thesis MSTR-2025-110

BibliographyChaudhari, Aishwarya: Visual Analysis of Image Data with Generative Models.
University of Stuttgart, Faculty of Computer Science, Electrical Engineering, and Information Technology, Master Thesis No. 110 (2025).
58 pages, english.
Abstract

The increasing ubiquity of high-resolution digital imagery across domains such as professional photography, curatorial practice, and documentation necessitates advanced tools for effective semantic organization and retrieval. This thesis presents the design, implementation, and evaluation of an interactive system that enables semantic filtering and visual exploration of large-scale image datasets using generative AI models, with particular emphasis on usability, transparency, and creative flexibility. Leveraging the capabilities of the GPT-4 language model, the system interprets natural language prompts to filter, cluster, and rank images based on semantic similarity. Central to the interaction paradigm is a visual semantic logic tree that externalizes user intent, facilitates prompt reuse, and supports iterative exploration. In addition, a projection-based clustering view allows for context-aware navigation through thematic image groupings, thereby enhancing discovery and interpretability. The system was developed using a modular architecture comprising a Python-based backend and an interactive D3.js frontend. The framework supports multiple query modes—including binary, rating-based, and projection-driven workflows—offering users differentiated levels of semantic control. Evaluation through domain-specific case studies (e.g., wedding photography, travel documentation, art exhibition curation) demonstrates that the system significantly reduces manual effort while improving cognitive alignment, traceability, and creative agency. This thesis contributes to the broader field of human–AI collaboration by showcasing how large-scale language models can be integrated into visual analytics interfaces that prioritize explainability, adaptability, and domain relevance. The concluding discussion addresses current limitations and outlines future research directions, including domain-specific fine-tuning, multimodal interaction support, and reinforcement-based adaptive learning mechanisms.

Department(s)University of Stuttgart, Institute of Visualisation and Interactive Systems, Visualisation and Interactive Systems
Superviser(s)Sedlmair, Prof. Michael; Satkunarajan, Jena; Koch, Dr. Steffen
Entry dateAugust 13, 2026
New Report   New Article   New Monograph   Institute   Computer Science