A Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning
| dc.contributor.author | Kahil, Hussain | |
| dc.contributor.author | Välisuo, Petri | |
| dc.contributor.author | Elmusrati, Mohammed | |
| dc.contributor.department | fi=Digital Economy|en=Digital Economy| | |
| dc.contributor.orcid | https://orcid.org/0000-0001-9304-6590 | |
| dc.date.accessioned | 2026-09-07T08:02:00Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | This paper proposes a Lyapunov-guided post-action shielding mechanism for deep reinforcement learning (DRL) controllers under bounded actuation disturbances. In addition, an energy-safety requirement is formulated as a one-step energy threshold constraint that keeps the predicted next-state energy proxy within a prescribed limit, using a bounded-disturbance worst-case check. Simulation results show that the proposed mechanism substantially reduces constraint violations under actuation noise while preserving the nominal policy behavior whenever possible. | en |
| dc.description.reviewstatus | fi=vertaisarvioitu|en=peerReviewed| | |
| dc.format.pagerange | 1649-1654 | |
| dc.identifier.citation | Kahil, H., Välisuo, P., & Elmusrati, M. (2026). A Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning. In 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT), 1649-1654. IEEE. https://doi.org/10.1109/CoDIT70676.2026.11630899 | |
| dc.identifier.isbn | 979-8-3195-2077-7 | |
| dc.identifier.uri | https://osuva.uwasa.fi/handle/11111/21269 | |
| dc.identifier.urn | URN:NBN:fi-fe20260907123261 | |
| dc.language.iso | en | |
| dc.publisher | IEEE | |
| dc.relation.conference | 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT) | |
| dc.relation.doi | https://doi.org/10.1109/codit70676.2026.11630899 | |
| dc.relation.isbn | 979-8-3195-2078-4 | |
| dc.relation.ispartof | 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT) | |
| dc.relation.ispartofjournal | International conference on control, decision and information technologies | |
| dc.relation.issn | 2576-3555 | |
| dc.relation.issn | 2576-3547 | |
| dc.relation.url | https://doi.org/10.1109/CoDIT70676.2026.11630899 | |
| dc.relation.url | https://urn.fi/URN:NBN:fi-fe20260907123261 | |
| dc.rights.copyright | © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. | |
| dc.source.identifier | 2-s2.0-105047837735 | |
| dc.source.identifier | c976c8f5-fb06-4bb3-8810-dd1804abc7c0 | |
| dc.source.metadata | SoleCRIS | |
| dc.subject | safe reinforcement learning | |
| dc.subject | Lyapunov stability | |
| dc.subject | post-action shielding | |
| dc.subject | energy threshold constraints | |
| dc.subject | learned controller robustness | |
| dc.subject.discipline | fi=Automaatiotekniikka|en=Automation Technology| | |
| dc.subject.discipline | fi=Tietoliikennetekniik|en=Telecommunications| | |
| dc.title | A Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning | |
| dc.type.okm | fi=A4 Vertaisarvioitu artikkeli konferenssijulkaisussa|en=A4 Article in conference proceedings (peer-reviewed)| | |
| dc.type.publication | article | |
| dc.type.version | acceptedVersion |
Tiedostot
1 - 1 / 1
