PL EN
Training machine learning agent behaviors in Unity: Proximal Policy Optimization vs. Soft Actor-Critic algorithms
 
More details
Hide details
1
Department of Computer Science, Lublin University of Technology
 
2
Lublin University of Technology
 
 
Publication date: 2026-08-31
 
 
Corresponding author
Maria Skublewska-Paszkowska   

Department of Computer Science, Lublin University of Technology
 
 
 
KEYWORDS
TOPICS
ABSTRACT
The gaming evolution is strictly associated with machine learning development, such as Reinforcement Learning. Thus, there is an ongoing need to study and apply machine learning algorithms to games to achieve human-like agent behavior. Among various algorithms, Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) are among the state-of-the-art reinforcement learning methods. They have been widely adopted due to their optimal balance of performance, stability, and ease of implementation. Factors such as the test environment, reward function, neural network architecture, and computational resource availability significantly influence the relative performance of both algorithms. Thus, the main aim of our study is to analyze PPO and SAC reinforcement learning methods across three Unity-created environments, using parallel learning with discrete and continuous actions. Using cumulative reward as the primary metric, SAC performed slightly better in simpler scenarios up to 1.1%, whereas PPO achieved better results in more complex environments up to 7.4%. The ablation study assesses hyperparameter sensitivity. Batch size had a stronger impact on SAC performance, whereas buffer size had a stronger impact on PPO.
Journals System - logo
Scroll to top