This thesis reports on the work carried out during my internship in the manufacturing company Siti B&T. During that period, the main activity I carried out was the development of an AI Agent that, through the WhatsApp messaging platform, communicates with employees to collect information and update the company's internal management system. The work was conducted along three principal directions: the development and deployment of the application for managing communication with employees and updating the management system; the analysis of the AI Agent's capabilities; and the training of the AI models to improve their performance. In order to concretely evaluate the AI's capabilities — without having prior data available — a custom synthetic dataset was created to simulate conversations with employees. These conversations were designed to resemble real-world scenarios as closely as possible, including cases in which the employee provides additional information in subsequent messages or communicates corrections to previously submitted data. In addition to enabling performance evaluation, this dataset also served as the training platform for improving the Agent. As the experimental foundation of this work, two training procedures were implemented for the two AI models used in the system. The first concerns the training of the Agent, for which an innovative dense reward mechanism was developed, this enables the model to learn from every action it performs during the conversation. Additionally, a reward component leveraging the semantic information from the optimal solution was implemented to guide the training process. The second training procedure targets a separate LLM used within the system to extract information and convert it into JSON format. In this case, the experiment explored different reward assignment mechanisms during training to determine which was most suitable. The thesis is structured as follows: starting from the definition of the problem, it proceeds to the design of the application and its deployment in the production environment. The theoretical foundations concerning the current state of Large Language Models (LLMs) and how they can be trained are then presented. Finally, the proposed training methodologies for the LLMs are described and the experimental results are discussed.
Reward Crumbs: Dense Reward Reinforcement Learning for Multi-Turn AI Agents in Conversational Information Extraction
ROSSI, NICOLÒ
2025/2026
Abstract
This thesis reports on the work carried out during my internship in the manufacturing company Siti B&T. During that period, the main activity I carried out was the development of an AI Agent that, through the WhatsApp messaging platform, communicates with employees to collect information and update the company's internal management system. The work was conducted along three principal directions: the development and deployment of the application for managing communication with employees and updating the management system; the analysis of the AI Agent's capabilities; and the training of the AI models to improve their performance. In order to concretely evaluate the AI's capabilities — without having prior data available — a custom synthetic dataset was created to simulate conversations with employees. These conversations were designed to resemble real-world scenarios as closely as possible, including cases in which the employee provides additional information in subsequent messages or communicates corrections to previously submitted data. In addition to enabling performance evaluation, this dataset also served as the training platform for improving the Agent. As the experimental foundation of this work, two training procedures were implemented for the two AI models used in the system. The first concerns the training of the Agent, for which an innovative dense reward mechanism was developed, this enables the model to learn from every action it performs during the conversation. Additionally, a reward component leveraging the semantic information from the optimal solution was implemented to guide the training process. The second training procedure targets a separate LLM used within the system to extract information and convert it into JSON format. In this case, the experiment explored different reward assignment mechanisms during training to determine which was most suitable. The thesis is structured as follows: starting from the definition of the problem, it proceeds to the design of the application and its deployment in the production environment. The theoretical foundations concerning the current state of Large Language Models (LLMs) and how they can be trained are then presented. Finally, the proposed training methodologies for the LLMs are described and the experimental results are discussed.| File | Dimensione | Formato | |
|---|---|---|---|
|
Rossi.Nicolo.pdf
Accesso riservato
Dimensione
16.14 MB
Formato
Adobe PDF
|
16.14 MB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14251/7543