This work presents the use of large language model (LLM) for human robot interaction. LLMs are increasingly being explored for human robot interaction to enable safe & effective communication between human and robots using natural language. This MSc thesis investigated the integration of LLM (Claude) with ROS2 using ROS MCP Sever, the work addresses the issues faced in safe and efficient human robot interaction using LLM. The system integrated Claude LLM to control robot using ROS MCP server. Using the tool calling capability of Claude it interacts with ROS MCP server, calling the functions defined in ROS MCP server, on the other hand ROS MCP server is used to access ROS2 topics through ROS bridge. In this architecture, under the umbrella of ROS MCP server two functions have been designed, namely robot task manager and safety watchdog. The purpose of robot task manager is to guide Claude in execution of a pick and place task. When user prompt LLM for pick and place task, Claude interacts with robot task manager in ROS MCP server , which defines a sequence of steps of be executed such as calculation of object grasp height , based on detected object height, move to pre pick position, move to grasp position, open gripper, close gripper, lift, move to transit position, move to pre place position, move to place position, open gripper, retreat to safe position. The safety watchdog function defines the workspace safety limits, joint limits, velocity and acceleration limits. In this project MoveIt planner is used for motion planning of UR10e robot. When Moveit Plans the trajectory each trajectory passes through safety watchdog layer, which samples the waypoint and check for safety violation, the trajectory is safe to executed only when it passes the safety check. When a trajectory generated by MoveIt planner fails in the safety check by watchdog layer, watchdog prompts Claude to replan, Claude then interacts with robot task manager which guides Claude to try with different perturbations of pose. The system also uses intel real sense depth camera, the depth data from camera combine with YOLO algorithm is used to detect the objects and update the planning scene. In this project, a probabilistic analysis with and without safety watch dog and robot task manager has been conducted using URSIM simulation, the analysis shows that when robot task manager and safety watchdog layer are used the success rate of Claude LLM executing task without safety violations increased to 76 % , while without robot task manager and safety watchdog layer the success rate of Claude LLM executing task without safety violation is only 25% .The analysis is conducted using 200 simulations, 100 simulation for each. To validate the framework with robot task manager and safety watchdog layer on real robot, UR10e robot has been used. The experimental setup consisted of UR10e robot, PC with Claude, ROS MCP server, ROS2, MoveIt planner, and RVIZ for visualization. For object detection, intel real sense depth camera has been calibrated using eye on base methods, and is used along with YOLO algorithm. The framework has been successfully validated in this experiment. The average task execution time observed during multiple experiments is around 4 minutes. The system performed quite slow; this is due to nature of LLM itself and re-planning attempts by MoveIt due to safety violations.

LLM for Human Robot Interaction LLM per l'interazione uomo-robot

UDDIN, SYED ZIA
2025/2026

Abstract

This work presents the use of large language model (LLM) for human robot interaction. LLMs are increasingly being explored for human robot interaction to enable safe & effective communication between human and robots using natural language. This MSc thesis investigated the integration of LLM (Claude) with ROS2 using ROS MCP Sever, the work addresses the issues faced in safe and efficient human robot interaction using LLM. The system integrated Claude LLM to control robot using ROS MCP server. Using the tool calling capability of Claude it interacts with ROS MCP server, calling the functions defined in ROS MCP server, on the other hand ROS MCP server is used to access ROS2 topics through ROS bridge. In this architecture, under the umbrella of ROS MCP server two functions have been designed, namely robot task manager and safety watchdog. The purpose of robot task manager is to guide Claude in execution of a pick and place task. When user prompt LLM for pick and place task, Claude interacts with robot task manager in ROS MCP server , which defines a sequence of steps of be executed such as calculation of object grasp height , based on detected object height, move to pre pick position, move to grasp position, open gripper, close gripper, lift, move to transit position, move to pre place position, move to place position, open gripper, retreat to safe position. The safety watchdog function defines the workspace safety limits, joint limits, velocity and acceleration limits. In this project MoveIt planner is used for motion planning of UR10e robot. When Moveit Plans the trajectory each trajectory passes through safety watchdog layer, which samples the waypoint and check for safety violation, the trajectory is safe to executed only when it passes the safety check. When a trajectory generated by MoveIt planner fails in the safety check by watchdog layer, watchdog prompts Claude to replan, Claude then interacts with robot task manager which guides Claude to try with different perturbations of pose. The system also uses intel real sense depth camera, the depth data from camera combine with YOLO algorithm is used to detect the objects and update the planning scene. In this project, a probabilistic analysis with and without safety watch dog and robot task manager has been conducted using URSIM simulation, the analysis shows that when robot task manager and safety watchdog layer are used the success rate of Claude LLM executing task without safety violations increased to 76 % , while without robot task manager and safety watchdog layer the success rate of Claude LLM executing task without safety violation is only 25% .The analysis is conducted using 200 simulations, 100 simulation for each. To validate the framework with robot task manager and safety watchdog layer on real robot, UR10e robot has been used. The experimental setup consisted of UR10e robot, PC with Claude, ROS MCP server, ROS2, MoveIt planner, and RVIZ for visualization. For object detection, intel real sense depth camera has been calibrated using eye on base methods, and is used along with YOLO algorithm. The framework has been successfully validated in this experiment. The average task execution time observed during multiple experiments is around 4 minutes. The system performed quite slow; this is due to nature of LLM itself and re-planning attempts by MoveIt due to safety violations.
2025
Large Language Model
HRI
ROS
MCP Server
Vision Perception
File in questo prodotto:
File Dimensione Formato  
UDDIN.SYEDZIA.pdf

Accesso riservato

Dimensione 1.06 MB
Formato Adobe PDF
1.06 MB Adobe PDF

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14251/6849