Tuning PI Gains Using Reinforcement Learning
R2026bProportional Integral (PI) tuning is the task of adjusting the proportional, integral and derivative gains of a PI controller to obtain stable, responsive and accurate closed-loop behavior. For relatively simple control tasks with few tunable parameters, model-based tuning techniques can get good results with a faster tuning process compared to model-free Reinforcement learning (RL)-based methods. See Control System Design and Tuning (Control System Toolbox) and Simulink Design Optimization for more information. However, for highly nonlinear or adaptive systems, you can get better results by using reinforcement learning methods to tune PI gains.
A more general setup for control applications is illustrated in Reinforcement Learning for Control Systems Applications.
The following sections explain how to use reinforcement learning in three PI tuning scenarios of increasing complexity.
Single Fixed Gains Across Multiple Operating Points

In this scenario, you know that the plant initial and operating conditions do not substantially change or they do not substantially affect the plant behavior. Here, you expect a single, fixed, PI controller to be able to drive the plant towards the desired behavior, independently on the plant initial condition. The operating conditions can also be called the context.
Therefore, the goal is to find one set of gains that performs well across different operating conditions.
To achieve this goal, treat the PI controller as a linear RL policy where the weights directly map to PI gains. Learning the best policy weights is equivalent to tuning the PI gains. The policy receives the error signal and outputs the control command.
During training, an RL environment step corresponds to a simulation step, and an RL episode corresponds to a simulation run. The RL agent that learns the policy is in feedback loop with the plant, and receives a reward corresponding to the current value of the error signal.
For an example that shows how to use reinforcement learning to tune the parameters of a controller, see Tune Fixed PI Gains Using Reinforcement Learning. You can also use this approach for parameter tuning or calibration applications.
Possible alternatives: For many PI tuning problems, model-based tuning algorithm (see Tune a Control System Using Control System Tuner (Simulink Control Design)), swarm particle optimization (see Field-Oriented Control of PMSM Using Fuzzy PI Controller (Fuzzy Logic Toolbox)), or Bayesian optimization can be simpler and more sample-efficient.
Fixed Set of Gains per Initial Condition

In this scenario, the plant can typically start from a wider range of initial condition. Also, the plant conditions substantially affect the plant behavior, but they do not change during plant operation. Here, you don't expect that a single set of gains can always work, but you think that, given the plant initial condition, there exist a corresponding fixed set of gains that can drive the plant towards the desired behavior.
Therefore, the goal is to learn a mapping from each possible initial condition to its PI gains, assuming that the operating conditions do not change while the plant is operating. Here, the RL policy associates one operating condition to one proportional and one integral gain.
To achieve this goal you use a contextual bandit approach (also referred to as single-step RL) to select the PI gains based on the current operating condition.
During training, each RL environment step corresponds to a single simulation run, during which the PI gains remain fixed. At the next training step, the RL environment changes the operating condition, the agent tries a new set of PI gains, and it receives a reward that depends on the error accumulated during the simulation. With this approach, specifying the range of the actions (PI gains) is recommended for effective training.
For an example that shows this approach, see Schedule PI Gains Per Initial Condition Using Reinforcement Learning.
For an example on contextual bandits, see Train Reinforcement Learning Agent for Simple Contextual Bandit Problem.
Note that, with this approach, you also have the option to periodically restart the policy to adapt the PI gains to a new initial condition, that is to a more recent operating condition. This might be appropriate when the plant conditions change slowly during operation.
Possible alternatives: More traditional approaches (see systune (Simulink Control Design) for
example) can also be a good choice. If your context space is small, or you do not know the
range of the PI gains, the first approach can be a good option. If the context changes
during an episode, try the third approach (online adaptive PI gain adjustment).
Dynamically Adapt Gains Online

This scenario is more general than the previous ones (that is it can include the previous scenarios as particular cases). In this scenario, the plant operating conditions affect the plant behavior and can also substantially change during operation. Here, you expect the controller to work only if the PI gains continuously adapt to the current operating conditions, during plant operation.
Therefore the goal is to continuously adapt the PI gains during plant operation to improve performance.
To achieve the goal you use an RL policy that updates the PI gains at each time step based on the current error. You can also use the context as additional observation.
This approach is similar to the first one, because it does not rely on a contextual bandit setting. As in the first approach, during training, an RL environment step corresponds to a simulation step, and an RL episode corresponds to a simulation run. The RL agent that learns the policy is in feedback loop with the plant, and receives a reward corresponding to the current value of the error signal.
However, differently from the first approach, the RL agent also observes a context, which (differently from the second approach) changes during plant operation. As a consequence, the policy does not consist of just two weights, (one proportional and one integral gain) but, similarly to the second approach, it is a mapping between both error and context and a set of PI gains.
For an example illustrating how to address this third scenario, see Dynamically Adapt PI Gains Online Using Reinforcement Learning.
Possible alternatives: Using an RL based controller, in which the output of the RL agent is the plant input instead of the PI gains, as in Reinforcement Learning for Control Systems Applications. Gain Scheduling (Control System Toolbox) or Adaptive and Time-Varying MPC (Model Predictive Control Toolbox) can also work for plants that have a dynamics which is reasonably well known. For an example that shows how to implement an RL-based controller, see Control Water Level in a Tank Using a DDPG Agent.
See Also
Functions
train|sim|rlSimulinkEnv