DEEP REINFORCEMENT LEARNING-BASED MAPLESS NAVIGATION ALGORITHM FOR HUSKY ROBOT WITH ADAPTIVE CURRICULUM LEARNING

Authors

  • Dmytro Petrenko Igor Sikorsky Kyiv Polytechnic Institute, faculty of robotics and instrumentation engineering, department of automation and non-destructive testing systems, Kyiv, Ukraine https://orcid.org/0009-0003-7670-555X
  • Anatolii Protasov Igor Sikorsky Kyiv Polytechnic Institute, faculty of robotics and instrumentation engineering, department of automation and non-destructive testing systems, Kyiv, Ukraine https://orcid.org/0000-0002-2965-3334

DOI:

https://doi.org/10.20535/kpisn.2026.2.357072

Keywords:

machine learning with reinforcement;, navigation without cartography; , mobile robots;, control systems; robot autonomy. , robot autonomy.

Abstract

Issues. Mobile wheeled robots are widely used today in internal logistics environments of various industries, in particular in warehouses, production shops and sorting centers. However, in such environments, the robot operates under conditions of high variability, where the configuration of aisles, the placement of loads and the presence of people can change in real time. Therefore, the development of autonomous navigation of the robot in complex and dynamic environments is becoming one of the key tasks of robotics.

Research objective. The method of this work increases the efficiency of navigation of a four-wheeled autonomous mobile robot without maps (mapless) in dynamic logistics environments by applying adaptive machine learning with reinforcement (Reinforcement Learning, RL) based on the curriculum.

Methods. To achieve the set goal and solve the tasks of autonomous navigation in logistics environments, a comprehensive methodological approach was used, based on the following methods:

- simulation method: Using the PyBullet physical engine to create a virtual logistics environment. This allowed for safe and high-repetition training, simulating the real-world kinematics of the Husky four-wheeled robot;

- deep Reinforcement Learning Methods: Application and comparison of three state-of-the-art architectures: SAC (Soft Actor-Critic), PPO (Proximal Policy Optimization), TD3 (Twin Delayed DDPG);

- adaptive Curriculum Learning (ACL): A step-by-step learning technique that automatically changes the complexity of the environment (number and speed of obstacles) depending on the current success of the agent. This avoids the degradation policy at early stages.

Results. The combination of adaptive learning (ACL) of the Husky mobile robot in complex dynamic environments and hybrid reward functions increased the navigation success rate by 35–40% compared to direct learning in complex scenarios. The conducted comparative analysis of machine learning algorithms demonstrated the highest adaptability to non-stationary conditions and the best efficiency of using experience is the Soft Actor-Critic (SAC) algorithm, which achieves success in 89% of cases of application with a minimum number of collisions.

Conclusions. A comprehensive approach to map-free navigation of the autonomous mobile robot Husky in complex dynamic environments is proposed, based on a combination of adaptive learning and a physically interpreted hybrid reward function. Experiments conducted in the PyBullet simulator confirmed that a stepwise and adaptive increase in the complexity of the environment stabilizes the learning process of modern deep reinforcement learning algorithms (PPO, SAC, and TD3).

References

D.V. Petrenko, A.H. Protasov, “Analysis of the effectiveness of Reinforcement Learning algorithms to increase the autonomy of mobile robots,” [Analiz efektyvnosti alhorytmiv Reinforcement Learning dlya pidvyshchennya avtonomnosti mobilnykh robotiv], zhurnal “Tekhnichna diahnostyka ta neruynivnyy kontrol”, 2025, №3, pp. 24-31. Retrieved from: https://doi.org/10.37434/tdnk2025.03.03

Wei, C., Li, Y., Ouyang, Y. et al. Deep Reinforcement Learning with Heuristic Corrections for UGV Navigation. J Intell Robot Syst 109, 18, 2023. Retrieved from: https://doi.org/10.1007/s10846-023-01950-y

M. Anca, J. D. Thomas, D. Pedamonti, M. Hansen та M. Studley. “Achieving Goals Using Reward Shaping and Curriculum Learning”, у Lecture Notes in Networks and Systems. Cham: Springer Nat. Switz., 2023, pp. 316–331. Retrieved from: https://doi.org/10.1007/978-3-031-47454-5_24

H. Xue, B. Hein, M. Bakr, G. Schildbach, B. Abel, E. Rueckert. Using Deep Reinforcement Learning with Automatic Curriculum Learning for Mapless Navigation in Intralogistics. Appl. Sci. 2022, 12, 3153. Retrieved from: https://doi.org/10.3390/app12063153

N. Wang, Y. Wang, Y. Zhao, Y. Wang та Z. Li “Sim-to-Real: Mapless Navigation for USVs Using Deep Reinforcement Learning”, J. Mar. Sci. Eng., т. 10, № 7, 895 p., 2022. Retrieved from: https://doi.org/10.3390/jmse10070895

“Husky A300 unmanned ground vehicle”. Clearpathrobotics. Accessed: 1 Feb. 2026. [Online]. Retrieved from: https://clearpathrobotics.com/husky-a300-unmanned-ground-vehicle-robot/

Downloads

Published

2026-06-29