Training machine learning agent behaviors in Unity: Proximal Policy Optimization vs. Soft Actor-Critic algorithms
Więcej
Ukryj
1
Department of Computer Science, Lublin University of Technology
2
Lublin University of Technology
Data publikacji: 31-08-2026
SŁOWA KLUCZOWE
DZIEDZINY
STRESZCZENIE
The gaming evolution is strictly associated with machine learning development, such as Reinforcement Learning. Thus, there is an ongoing need to study and apply machine learning algorithms to games to achieve human-like agent behavior. Among various algorithms, Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) are among the state-of-the-art reinforcement learning methods. They have been widely adopted due to their optimal balance of performance, stability, and ease of implementation. Factors such as the test environment, reward function, neural network architecture, and computational resource availability significantly influence the relative performance of both algorithms. Thus, the main aim of our study is to analyze PPO and SAC reinforcement learning methods across three Unity-created environments, using parallel learning with discrete and continuous actions. Using cumulative reward as the primary metric, SAC performed slightly better in simpler scenarios up to 1.1%, whereas PPO achieved better results in more complex environments up to 7.4%. The ablation study assesses hyperparameter sensitivity. Batch size had a stronger impact on SAC performance, whereas buffer size had a stronger impact on PPO.