| Bibliography | Chaudhari, Aishwarya: Visual Analysis of Image Data with Generative Models. University of Stuttgart, Faculty of Computer Science, Electrical Engineering, and Information Technology, Master Thesis No. 110 (2025). 58 pages, english.
|
| Abstract | The increasing ubiquity of high-resolution digital imagery across domains such as professional photography, curatorial practice, and documentation necessitates advanced tools for effective semantic organization and retrieval. This thesis presents the design, implementation, and evaluation of an interactive system that enables semantic filtering and visual exploration of large-scale image datasets using generative AI models, with particular emphasis on usability, transparency, and creative flexibility. Leveraging the capabilities of the GPT-4 language model, the system interprets natural language prompts to filter, cluster, and rank images based on semantic similarity. Central to the interaction paradigm is a visual semantic logic tree that externalizes user intent, facilitates prompt reuse, and supports iterative exploration. In addition, a projection-based clustering view allows for context-aware navigation through thematic image groupings, thereby enhancing discovery and interpretability. The system was developed using a modular architecture comprising a Python-based backend and an interactive D3.js frontend. The framework supports multiple query modes—including binary, rating-based, and projection-driven workflows—offering users differentiated levels of semantic control. Evaluation through domain-specific case studies (e.g., wedding photography, travel documentation, art exhibition curation) demonstrates that the system significantly reduces manual effort while improving cognitive alignment, traceability, and creative agency. This thesis contributes to the broader field of human–AI collaboration by showcasing how large-scale language models can be integrated into visual analytics interfaces that prioritize explainability, adaptability, and domain relevance. The concluding discussion addresses current limitations and outlines future research directions, including domain-specific fine-tuning, multimodal interaction support, and reinforcement-based adaptive learning mechanisms.
|