Master Thesis MSTR-2025-66

BibliographyHengst, Nora Marie: An Analysis of Chatbot Conversations in AI-Assisted Software Engineering.
University of Stuttgart, Faculty of Computer Science, Electrical Engineering, and Information Technology, Master Thesis No. 66 (2025).
87 pages, english.
Abstract

Introduction: The rapid evolution of Large Language Model (LLM)-based chatbot capabilities has not been matched by a corresponding development in how users, particularly software engineers, communicate with them. Meaning, that there is little research on best practices on how to interact goal driven with an LLM-based chatbot. Software engineers often lack the necessary knowledge to formulate targeted queries for LLM-based chatbots in order to obtain the desired output. Focusing on the user side, software engineers often don’t have a lot of training in using Artificial Intelligence (AI) tools like LLM-based chatbots. This often leads to inefficient interactions, frustration, and suboptimal results. Objective: This work addresses this gap by developing a guideline for more effective and efficient communication between software engineers and LLM-based chatbots in a software engineering context. Methodology: The methodology consists of a mixed-methods approach containing automated and manual analysis. A foundational literature review on handling incorrect AI responses in user to LLM-based chatbot conversations in general was carried out. It also targeted, how to carry out conversation analysis methods. After that, an existing dataset of approximately 90,000 historical chat conversations was used and filtered down with the use of keywords to approximately 8,000 relevant English-language dialogues with explicit software engineering content. Automated labeling was employed to categorize conversation topics, while sentiment and success-failure analyses were conducted. Subsequently, a detailed manual analysis of 40 conversations was conducted to identify recurring patterns and user strategies. Results: The main results reveal that most conversations fall into the categories of Setup, Knowledge Sharing & Learning, Code Development, and Research & Writing. While automated labeling was successful for topic classification, sentiment and success-failure analyses proved largely ineffective due to a lack of clear indicators within the chats. The manual analysis, however, identified significant regularities in user behavior, leading to the extraction of four central themes that occurred at least eight times in the analyzed conversations. These central themes were: Questioning, Instruction, Scenario, and Instruction on Output Format. Meaning that the named categories refer to how a user has phrased the query to the LLM-based chatbot. Based on these findings, a set of guidelines was created to assist users in formulating prompts and provide troubleshooting steps for unsuccessful conversations. Conclusion: In conclusion, this work successfully identifies common patterns in how software engineers interact with LLM-based chatbots. It also provides a foundational framework for improving these interactions. The results furthermore highlight the challenges of an automated analysis for evaluating chat success and sentiment, suggesting a need for future research to focus specifically on failed conversations to better distinguish them from successful ones. Moreover, it opens the question on what the indicators and a definition of a failed or a successful conversation are. Furthermore, the proposed guidelines require additional empirical testing to validate their applicability and effectiveness in real-world scenarios.

Department(s)University of Stuttgart, Institute of Software Technology, Empirical Software Engineering
Superviser(s)Wagner, Prof. Stefan; Munoz Baron, Marvin
Entry dateDecember 19, 2025
   Publ. Institute   Publ. Computer Science