PPO algorithm training problem in Reinforcement Learning Toolbox

Question

DQ LEE am 28 Jun. 2023

0
Verknüpfen

Direkter Link zu dieser Frage

https://de.mathworks.com/matlabcentral/answers/1989158-ppo-algorithm-training-problem-in-reinforcement-learning-toolbox

Kommentiert: 轩 am 31 Dez. 2023

In the PPO training algorithm , here mentioned “For each experience sequence that does not contain a terminal state, N is equal to the ExperienceHorizon option value. Otherwise, N is less than ExperienceHorizon and SN is the terminal state.” ,

Here's my question :When N is smaller than ExperienceHorizon and N is also smaller than the size of mini-batch data, and this continues for multiple consecutive episodes, When does the algorithm update the parameters in this case?

AND another one question is :When will the PPO parameter be updated under the following parameter Settings:

agentOpts = rlPPOAgentOptions(...

'ExperienceHorizon',10000,...

'MiniBatchSize',64,...

'NumEpoch',3,...)

trainOpts = rlTrainingOptions(...

'MaxEpisodes',10000,...

'MaxStepsPerEpisode',30,... )

0 Kommentare
-2 ältere Kommentare anzeigen-2 ältere Kommentare ausblenden

Melden Sie sich an, um zu kommentieren.

Melden Sie sich an, um diese Frage zu beantworten.

Answer 1

Takeshi Takahashi am 5 Jul. 2023

0
Verknüpfen

Direkter Link zu dieser Antwort

https://de.mathworks.com/matlabcentral/answers/1989158-ppo-algorithm-training-problem-in-reinforcement-learning-toolbox#answer_1267773

When N is smaller than ExperienceHorizon and N is also smaller than MiniBatchSize, the PPO agent uses N experiences to update its parameters at the end of the episode.

So, if MaxStepsPerEpisode = 30, ExperienceHorizon = 10000, and MiniBatchSize is 64, the PPO agent uses 30 or fewer experiences (when the episode terminates early) to update its parameters at the end of each episode.

2 Kommentare
Keine anzeigenKeine ausblenden

轩 am 31 Dez. 2023

So who deside the value N when the episode does not be stopped by reaching ExperienceHorizon and terminal state ?

Thank you for your explanation in advace.

轩 am 31 Dez. 2023

Maybe I have found the answer in the document Create Policies and Value Functions - MATLAB & Simulink - MathWorks Benelux

"When using PG agents, the learning trajectory length (that is the sequence of input data that the network uses for learning) for the RNN is the whole episode. For an AC agent, the NumStepsToLookAhead property of its options object is treated as the training trajectory length (except when training in parallel, in which case NumStepsToLookAhead is ignored and the whole episode is used as trajectory length). For a PPO agent, the trajectory length is the MiniBatchSize property of its options object."

Melden Sie sich an, um zu kommentieren.

PPO algorithm training problem in Reinforcement Learning Toolbox

0 Kommentare
-2 ältere Kommentare anzeigen-2 ältere Kommentare ausblenden

Akzeptierte Antwort

2 Kommentare
Keine anzeigenKeine ausblenden

Weitere Antworten (0)

Siehe auch

Kategorien

Tags

Produkte

Community Treasure Hunt

PPO algorithm training problem in Reinforcement Learning Toolbox

0 Kommentare -2 ältere Kommentare anzeigen-2 ältere Kommentare ausblenden

Akzeptierte Antwort

2 Kommentare Keine anzeigenKeine ausblenden

Weitere Antworten (0)

Siehe auch

Kategorien

Tags

Produkte

Community Treasure Hunt

0 Kommentare
-2 ältere Kommentare anzeigen-2 ältere Kommentare ausblenden

2 Kommentare
Keine anzeigenKeine ausblenden