Research

OUC Made New Progress in Embodied Intelligence for Robotics

Recently, a research group led by Associate Professor Li Guangliang of the Faculty of Information Science and Engineering at Ocean University of China (OUC), in collaboration with Professor Xue Jianru’s team at the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence at Xi’an Jiaotong University, made new progress in embodied intelligence for robotics. The findings were published in an article entitled “Generative Adversarial Self-Imitation Learning with Large Language Model Feedback for Robot Control and Navigation” in IEEE Transactions on Robotics. 


In recent years, reinforcement learning and imitation learning have advanced significantly in robot control and autonomous navigation. In complex robotic tasks, however, robots’ capacity for autonomous learning is directly constrained by the effectiveness of reward-function design and the quality of expert demonstrations. To address this challenge, the team proposed a novel generative adversarial self-imitation learning method incorporating large language model (LLM) feedback: Generative Adversarial Self-Imitation Learning from Demonstration and Large Language Model Feedback (GASL³MF). The method draws on the rich commonsense knowledge and reasoning capabilities encoded in LLMs to evaluate robot behavior. A feedback model trained on these evaluations then replaces frequent calls to the LLM, reducing computational costs while guiding the robot to continually refine its policy.

 


Unlike conventional imitation-learning methods that rely on high-quality expert demonstrations, GASL³MF can learn from low-quality or even failed demonstrations. As the robot performs a task, it continuously generates new trajectories and uses an LLM feedback model to assess their quality. Trajectories that outperform the original demonstrations are progressively added to the demonstration buffer, replacing lower-quality data and enabling the robot’s continuous self-optimization and self-learning. This mechanism overcomes a key limitation of conventional methods, which can imitate demonstrations but struggle to outperform them, and allows the robot to gradually learn a near-optimal control policy. The work further broadens the application of LLMs in robot learning and offers new technical approaches to autonomous robot control in complex environments, intelligent navigation, human-robot collaboration, and underwater robotics. It is of considerable significance for advancing key technologies in embodied intelligence and intelligent robotics.