Bachelor Thesis BCLR-2025-118

BibliographyLippmann, Tom: A Study of Programming-Related Conversations between Humans and LLMs.
University of Stuttgart, Faculty of Computer Science, Electrical Engineering, and Information Technology, Bachelor Thesis No. 118 (2025).
73 pages, english.
Abstract

Large Language Models (LLMs) have become integral tools for software development and everyday use, with millions of users engaging daily. Yet little is known about how users actually interact with them in programming-related contexts. To help close this research gap, this thesis investigates the structure, content, and coherence of programming-related conversations between humans and LLMs. We analyze two large-scale datasets, the LMSYS-Chat-1M and DevGPT, and apply a keyword-based filtering approach to collect 120,982 and 2,152 programming-related conversations, respectively. Using a combination of quantitative and qualitative methods, we analyze message and conversation length, programming language distribution, common topics and tasks, prompt specificity, task difficulty, semantic coherence, and failing conversations. Our research reveals the following: Programming-related conversations are typically short, with most user messages consisting of two to three sentences and most conversations containing two turns or less. LLM-generated code snippets are predominantly written in Python. We identify ten distinct, commonly occurring topics in programming-related conversations through manual annotation of user messages and clustering. The most commonly appearing topic is code generation. More than 90% of programming-related conversations are coherent, and between 90.3-97.4% end on the same or a similar topic as they started with. When only multi-turn conversations are examined, the number of incoherent conversations ranges from 11.9 to 34.8%. We find indicators for prompt specificity being roughly uniformly distributed, while task difficulty primarily depends on the user group. Developers engage in more complex tasks, while casual users discuss simpler ones. Our analysis of unsuccessful conversations indicates that most users do not provide any feedback when receiving unhelpful responses. Based on these findings, we recommend improvements to LLM design, including self-verification mechanisms, models tailored to the user’s programming experience, and features that encourage users to participate more actively in conversations and provide feedback.

Department(s)University of Stuttgart, Institute of Software Technology, Software Lab - Program Analysis
Superviser(s)Pradel, Prof. Michael; Eghbali, Aryaz
Entry dateJune 10, 2026
   Publ. Institute   Publ. Computer Science