A

Safe Value Functions

IEEE Transactions on Automatic Control, vol. 68, pp. 2743–2757

Abstract

Safety constraints and optimality are important but sometimes conflicting criteria for controllers. Although these criteria are often solved separately with different tools to maintain formal guarantees, it is also common practice in reinforcement learning (RL) to simply modify reward functions by penalizing failures, with the penalty treated as a mere heuristic. We rigorously examine the relationship of both safety and optimality to penalties, and formalize sufficient conditions forsafe value functions (SVFs):value functions that are both optimal for a given task, and enforce safety constraints. We reveal this structure by examining when rewards preserve viability under optimal control, and show that there always exists a finite penalty that induces an SVF. This penalty is not unique, but upper-unbounded: larger penalties do not harm optimality. Although it is often not possible to compute the minimum required penalty, we reveal clear structure of how the penalty, rewards, discount factor, and dynamics interact. This insight suggests practical, theory-guided heuristics to design reward functions for control problems where safety is important.

Authors 4

  1. RWTH Aachen University

    Affiliation as printed

    Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany

    RWTH Aachen University

  2. Steve Heim Aachen

    RWTH Aachen University · Massachusetts Institute of Technology

    Affiliation as printed

    Biomimetic Robotics Lab, Massachusetts Institute of Technology, Cambridge, MA, USA

    RWTH Aachen University

  3. RWTH Aachen University

    Affiliation as printed

    Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany

  4. RWTH Aachen University · Max Planck Society

    Affiliation as printed

    Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany

    Max-Planck-Society

Cited by 6 stored of 6

6 results

No patents citing this paper on Lens.org (checked 2026-10-06).

References 86