A Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning

dc.contributor.authorKahil, Hussain
dc.contributor.authorVälisuo, Petri
dc.contributor.authorElmusrati, Mohammed
dc.contributor.departmentfi=Digital Economy|en=Digital Economy|
dc.contributor.orcidhttps://orcid.org/0000-0001-9304-6590
dc.date.accessioned2026-09-07T08:02:00Z
dc.date.issued2026
dc.description.abstractThis paper proposes a Lyapunov-guided post-action shielding mechanism for deep reinforcement learning (DRL) controllers under bounded actuation disturbances. In addition, an energy-safety requirement is formulated as a one-step energy threshold constraint that keeps the predicted next-state energy proxy within a prescribed limit, using a bounded-disturbance worst-case check. Simulation results show that the proposed mechanism substantially reduces constraint violations under actuation noise while preserving the nominal policy behavior whenever possible.en
dc.description.reviewstatusfi=vertaisarvioitu|en=peerReviewed|
dc.format.pagerange1649-1654
dc.identifier.citationKahil, H., Välisuo, P., & Elmusrati, M. (2026). A Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning. In 2026 12th International Conference on Control, Decision and Information Technologies (CoDIT), 1649-1654. IEEE. https://doi.org/10.1109/CoDIT70676.2026.11630899
dc.identifier.isbn979-8-3195-2077-7
dc.identifier.urihttps://osuva.uwasa.fi/handle/11111/21269
dc.identifier.urnURN:NBN:fi-fe20260907123261
dc.language.isoen
dc.publisherIEEE
dc.relation.conference2026 12th International Conference on Control, Decision and Information Technologies (CoDIT)
dc.relation.doihttps://doi.org/10.1109/codit70676.2026.11630899
dc.relation.isbn979-8-3195-2078-4
dc.relation.ispartof2026 12th International Conference on Control, Decision and Information Technologies (CoDIT)
dc.relation.ispartofjournalInternational conference on control, decision and information technologies
dc.relation.issn2576-3555
dc.relation.issn2576-3547
dc.relation.urlhttps://doi.org/10.1109/CoDIT70676.2026.11630899
dc.relation.urlhttps://urn.fi/URN:NBN:fi-fe20260907123261
dc.rights.copyright© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
dc.source.identifier2-s2.0-105047837735
dc.source.identifierc976c8f5-fb06-4bb3-8810-dd1804abc7c0
dc.source.metadataSoleCRIS
dc.subjectsafe reinforcement learning
dc.subjectLyapunov stability
dc.subjectpost-action shielding
dc.subjectenergy threshold constraints
dc.subjectlearned controller robustness
dc.subject.disciplinefi=Automaatiotekniikka|en=Automation Technology|
dc.subject.disciplinefi=Tietoliikennetekniik|en=Telecommunications|
dc.titleA Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning
dc.type.okmfi=A4 Vertaisarvioitu artikkeli konferenssijulkaisussa|en=A4 Article in conference proceedings (peer-reviewed)|
dc.type.publicationarticle
dc.type.versionacceptedVersion

Tiedostot

Näytetään 1 - 1 / 1
Ladataan...
Name:
nbnfi-fe20260907123261.pdf
Size:
2.92 MB
Format:
Adobe Portable Document Format

Kokoelmat