Research Article

A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability

Volume: 12 Number: 3 September 30, 2026

A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability

Abstract

This study introduces the Proposed Optimization Algorithm (POA), a swarmbased hybrid combining the Gorilla Troops Optimizer (GTO) and the Artificial Bee Colony (ABC) algorithm, to enhance reward maximization in control tasks. Neural networks were trained for both simple (pendulum) and complex (bipedal walker) environments. The POA algorithm was run 10 times, and based on the resulting median values, the best reward scores achieved were −117.771 in the pendulum environment and 30.936 in the bipedal walker environment. These reward values indicate a 0.244 improvement for the pendulum environment and a 39.569 improvement for the bipedal walker compared to the closest competitors (GTO). While there was no statistically significant difference between GTO and POA in the pendulum task, POA performed significantly better than all other algorithms in the bipedal walker environment (p < 0.05). To address the “black box” nature of reinforcement learning, the study integrated Shapley Value Theory for post-training analysis. This explainable AI (XAI) approach identified angular velocity as the primary driver of torque in the pendulum task and quantified the importance of observation parameters for the bipedal walker. The results provide both a high-performing optimization framework and a robust method for interpreting neural network decision-making in robotic control systems.

Keywords

Ethical Statement

No approval from the Board of Ethics is required.

References

  1. M. L. Puterman, Markov decision processes: Discrete stochastic dynamic programming, John Wiley & Sons, 2014.
  2. R. Ozalp, A. Ucar, C. Guzelis, Advancements in deep reinforcement learning and inverse reinforcement learning for robotic manipulation: Toward trustworthy, interpretable, and explainable artificial intelligence, IEEE Access 12 (2024) 51840–51858.
  3. N. Justesen, P. Bontrager, J. Togelius, S. Risi, Deep learning for video game playing, IEEE Transactions on Games 12 (1) (2020) 1–20.
  4. A. A. Abdellatif, N. Mhaisen, A. Mohamed, A. Erbad, M. Guizani, Reinforcement learning for intelligent healthcare systems: A review of challenges, applications, and open research issues, IEEE Internet of Things Journal 10 (24) (2023) 21982–22007.
  5. B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. Al Sallab, S. Yogamani, P. PÅLerez, Deep reinforcement learning for autonomous driving: A survey, IEEE Transactions on Intelligent Transportation Systems 23 (6) (2022) 4909–4926.
  6. I. Saleh, N. Borhan, A. Yunus, W. Rahiman, Comprehensive technical review of recent bio-inspired population-based optimization (BPO) algorithms for mobile robot path planning, IEEE Access 12 (2024) 20942–20961.
  7. L. Deng, S. Liu, Advancing photovoltaic system design: An enhanced social learning swarm optimizer with guaranteed stability, Computers in Industry 164 (2025) 104209.
  8. S. Xing, Z. Shao, W. Shao, J. Chen, D. Pi, Joint scheduling of hybrid flow-shop with limited automatic guided vehicles: A hierarchical learning-based swarm optimizer, Computers & Industrial Engineering 198 (2024) 110686.

Details

Primary Language

English

Subjects

Control Theoryand Applications

Journal Section

Research Article

Publication Date

September 30, 2026

Submission Date

May 11, 2026

Acceptance Date

July 23, 2026

Published in Issue

Year 2026 Volume: 12 Number: 3

APA
Bingöl, M. C. (2026). A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability. Journal of Advanced Research in Natural and Applied Sciences, 12(3), 264-281. https://doi.org/10.28979/jarnas.1949124
AMA
1.Bingöl MC. A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability. JARNAS. 2026;12(3):264-281. doi:10.28979/jarnas.1949124
Chicago
Bingöl, Mustafa Can. 2026. “A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems With Shapley Value Interpretability”. Journal of Advanced Research in Natural and Applied Sciences 12 (3): 264-81. https://doi.org/10.28979/jarnas.1949124.
EndNote
Bingöl MC (September 1, 2026) A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability. Journal of Advanced Research in Natural and Applied Sciences 12 3 264–281.
IEEE
[1]M. C. Bingöl, “A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability”, JARNAS, vol. 12, no. 3, pp. 264–281, Sept. 2026, doi: 10.28979/jarnas.1949124.
ISNAD
Bingöl, Mustafa Can. “A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems With Shapley Value Interpretability”. Journal of Advanced Research in Natural and Applied Sciences 12/3 (September 1, 2026): 264-281. https://doi.org/10.28979/jarnas.1949124.
JAMA
1.Bingöl MC. A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability. JARNAS. 2026;12:264–281.
MLA
Bingöl, Mustafa Can. “A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems With Shapley Value Interpretability”. Journal of Advanced Research in Natural and Applied Sciences, vol. 12, no. 3, Sept. 2026, pp. 264-81, doi:10.28979/jarnas.1949124.
Vancouver
1.Mustafa Can Bingöl. A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability. JARNAS. 2026 Sep. 1;12(3):264-81. doi:10.28979/jarnas.1949124

 

 

 

TR Dizin 20466
 

 

SAO/NASA Astrophysics Data System (ADS)    34270

                                                   American Chemical Society-Chemical Abstracts Service CAS    34922 

 

DOAJ 32869

EBSCO 32870

Scilit 30371                        

SOBİAD 20460

 

29804 JARNAS is licensed under a Creative Commons Attribution-NonCommercial 4.0 International Licence (CC BY-NC).