نشریه مهندسی مکانیک امیرکبیر

نشریه مهندسی مکانیک امیرکبیر

طراحی کنترل مقاوم وضعیت ماهواره سه‌درجه‌آزادی با تلفیق الگوریتم فوق‌پیچشی و یادگیری تقویتی عمیق و تنظیم فراپارامترها با طرح‌ریزی آزمایش تاگوچی

نوع مقاله : مقاله پژوهشی

نویسندگان
گروه مهندسی هوافضا، دانشکده فنی و مهندسی، دانشگاه اصفهان، اصفهان، ایران
چکیده
در این پژوهش، یک چارچوب کنترلی ترکیبی برای کنترل وضعیت ماهواره سه‌درجه‌آزادی در حضور نامعینی‌های پارامتری، اغتشاش‌های خارجی، محدودیت‌های عملگر و خطاهای اجرایی ارائه می‌شود. هسته‌ی مقاوم سامانه بر پایه الگوریتم فوق‌پیچشی طراحی شده است تا ضمن حفظ پایداری و روباستی، پدیده چترینگ کاهش یابد. به‌منظور بهبود دقت رهگیری و افزایش قابلیت سازگاری در مواجهه با غیرخطی‌بودن و عدم‌قطعیت‌های پیش‌بینی‌ناپذیر، از یادگیری تقویتی عمیق به‌عنوان یک مؤلفه اصلاحی تطبیقی در کنار کنترل‌کننده مقاوم استفاده شده است. در این راستا، سه الگوریتم شاخص شامل گرادیان سیاست تعیینی عمیق، گرادیان سیاست تعیینی عمیق تأخیردار دوقلو و بهینه‌سازی سیاست مجاورتی مورد مطالعه و مقایسه قرار گرفته‌اند. همچنین، برای کاهش هزینه محاسباتی تنظیم فراپارامترها و ایجاد یک رویه نظام‌مند و تکرارپذیر، از روش طرح‌ریزی آزمایش تاگوچی جهت بهینه‌سازی چندهدفه فراپارامترهای یادگیری تقویتی عمیق بهره گرفته شده است. معیار ارزیابی عملکرد شامل ترکیبی از دقت رهگیری زمان‌وزن‌شده و میزان تلاش کنترلی در نظر گرفته شده است. نتایج شبیه‌سازی عددی و پیاده‌سازی آزمایشگاهی روی شبیه‌ساز وضعیت ماهواره نشان می‌دهد ساختار ترکیبی پیشنهادی، در مقایسه با کنترل مقاوم پایه، می‌تواند زمان نشست و تلاش کنترلی را کاهش داده و استحکام در برابر اغتشاشات را بهبود دهد، در حالی که پایداری و دقت رهگیری حفظ می‌شود.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Robust Attitude Control of a Three-Degree-of-Freedom Satellite via Integration of the Super-Twisting Algorithm and Deep Reinforcement Learning with Hyperparameter Tuning Using Taguchi Design of Experiments

نویسندگان English

Mostafa Sarjoughian
Hojat Taei
Department of Aerospace Engineering, Faculty of Engineering, University of Isfahan, Isfahan, Iran
چکیده English

This study presents a hybrid control framework for the attitude regulation of a three-degree-of-freedom satellite subject to parametric uncertainties, external disturbances, actuator constraints, and implementation imperfections. The core robust controller is formulated using the Super-Twisting Algorithm, which guarantees finite-time convergence and robustness while effectively suppressing the high-frequency chattering typically associated with conventional sliding mode control. To enhance tracking precision and improve adaptability under nonlinear and uncertain conditions, deep reinforcement learning is incorporated as an adaptive compensator within the control loop. Three representative algorithms, namely Deep Deterministic Policy Gradient, Twin Delayed Deep Deterministic Policy Gradient, and Proximal Policy Optimization, are investigated and comparatively evaluated in terms of stability, convergence behavior, and control efficiency. To systematically tune the learning hyperparameters and reduce the computational burden associated with manual trial-and-error procedures, the Taguchi design of experiments method is employed to perform multi-objective optimization considering both tracking performance and control effort. The performance index is defined as a composite measure that combines time-weighted tracking error and control energy. Numerical simulations together with experimental validation on a satellite attitude simulator demonstrate that the proposed hybrid control architecture reduces settling time and control effort while improving disturbance rejection capability, without compromising stability or steady-state tracking accuracy.

کلیدواژه‌ها English

Satellite Attitude Control
Super-Twisting Algorithm
Deep Reinforcement Learning
Taguchi Design of Experiments
[1] K. Lu, Y. Xia, Finite‐time attitude control for rigid spacecraft‐based on adaptive super‐twisting algorithm, IET Control Theory & Applications, 8(15) (2014) 1465–1477.
[2] Y. Su, S. Shen, Adaptive predefined-time fault-tolerant attitude tracking control for rigid spacecraft with guaranteed performance, Acta Astronautica, 214 (2024) 677–688.
[3] C. Xiao, Y. Guo, C.-q. Xie, A.-j. Li, C.-q. Wang, Adaptive super-twisting sliding mode attitude coordination control for spacecraft formation flying with actuator saturation, Advances in Space Research, 72(10) (2023) 4244–4255.
[4] M. Khodaverdian, M. Malekzadeh, Fault-tolerant model predictive sliding mode control with fixed-time attitude stabilization and vibration suppression of flexible spacecraft, Aerospace Science and Technology, 139 (2023) 108381.
[5] X.-N. Shi, W. Chen, R. Li, Z.-G. Zhou, K. Wen, Prescribed Performance Attitude Tracking Control for Spacecraft under Multi-Constraint, in:  2020 39th Chinese Control Conference (CCC), IEEE, 2020, pp. 270–275.
[6] M. Tipaldi, R. Iervolino, P.R. Massenio, Reinforcement learning in spacecraft control applications: Advances, prospects, and challenges, Annual Reviews in Control, 54 (2022) 1–23.
[7] W. Retagne, J. Dauer, G. Waxenegger-Wilfing, Adaptive satellite attitude control for varying masses using deep reinforcement learning, Frontiers in Robotics and AI, 11 (2024) 1402846.
[8] K.-H. Lee, S. Lim, D.-H. Cho, H.-D. Kim, Development of fault detection and identification algorithm using deep learning for nanosatellite attitude control system, International Journal of Aeronautical and Space Sciences, 21(2) (2020) 576–585.
[9] S. Oghim, J. Park, H. Bang, H. Leeghim, Deep reinforcement learning-based attitude control for spacecraft using control moment gyros, Advances in Space Research, 75(1) (2025) 1129–1144.
[10] M. Wu, K. Guo, X. Li, Z. Lin, Y. Wu, T.A. Tsiftsis, H. Song, Deep reinforcement learning-based energy efficiency optimization for RIS-aided integrated satellite-aerial-terrestrial relay networks, IEEE Transactions on Communications, 72(7) (2024) 4163–4178.
[11] N.A. Mosali, S.S. Shamsudin, O. Alfandi, R. Omar, N. Al-Fadhali, Twin delayed deep deterministic policy gradient-based target tracking for unmanned aerial vehicle with achievement rewarding and multistage training, IEEE Access, 10 (2022) 23545–23559.
[12] M. Ran, J. Li, L. Xie, Reinforcement-learning-based disturbance rejection control for uncertain nonlinear systems, IEEE Transactions on Cybernetics, 52(9) (2021) 9621–9633.
[13] X. Zhang, X. Chen, L. Yao, C. Ge, M. Dong, Deep neural network hyperparameter optimization with orthogonal array tuning, in:  International conference on neural information processing, Springer, 2019, pp. 287–295.
[14] J.B. Chandar, M. Sivakumar, N. Lenin, R. Čep, S. Salunkhe, E.A. Nasr, C. Rathinasuriyan, A novel predictive model for abrasive waterjet deep hole drilling on AL7075 T6 using machine learning and evolutionary algorithmic approach, Scientific Reports, 15(1) (2025) 43951.
[15] N. Ranković, D. Ranković, GOAT method: Green Orthogonal Array Tuning method, Alexandria Engineering Journal, 133 (2025) 13–41.
[16] P. Arevalo, A. Cano, O. Fedoseienko, F. Jurado, A data-driven approach to microgrid fault detection and classification using Taguchi-optimized CNNs and wavelet transform, Applied Soft Computing, 170 (2025) 112667.
[17] L. Grbcic, M. Park, J. Müller, V. Zorba, W.A. de Jong, Artificial intelligence driven laser parameter search: Inverse design of photonic surfaces using greedy surrogate-based optimization, Engineering Applications of Artificial Intelligence, 143 (2025) 109971.
[18] C.-J. Lin, S.-Y. Jeng, C.-L. Lee, Hyperparameter Optimization of Deep Learning Networks for Classification of Breast Histopathology Images, Sensors & Materials, 33 (2021).
[19] M. Sarjoughian, M. Malekzadeh, N. Sayyaf, Hybrid Control of Spacecraft: Super-Twisting Algorithm Based on Taguchi-Driven Deep Reinforcement Learning, Results in Engineering,  (2026) 110530.
[20] S. Jamshidi, M. Mirzaei, M. Malekzadeh, Applied optimal control of spacecraft simulator subject to failures of reaction wheels, Arabian Journal for Science and Engineering, 49(2) (2024) 1697–1712.
[21] R.S. Sutton, A.G. Barto, Reinforcement learning: An introduction, MIT press Cambridge, 1998.
[22] S. Fujimoto, H. Hoof, D. Meger, Addressing function approximation error in actor-critic methods, in:  International conference on machine learning, PMLR, 2018, pp. 1587–1596.
[23] A. Surriani, O. Wahyunggoro, A trajectory control for bipedal walking robot using stochastic-based continuous deep reinforcement learning,  (2023).
[24] S. Qi, L. Lu, F. Ziruo, B. Xingzi, L. Huaqiu, C. Wen, Y. Jinpei, Efficient and fair PPO-based integrated scheduling method for multiple tasks of SATech-01 satellite, Chinese Journal of Aeronautics, 37(2) (2024) 417–430.