Oscillation of Episode Q0 during DDPG training

2 Ansichten (letzte 30 Tage)

Ältere Kommentare anzeigen

Heesu Kim am 6 Apr. 2021

0
Verknüpfen

Direkter Link zu dieser Frage

https://de.mathworks.com/matlabcentral/answers/794607-oscillation-of-episode-q0-during-ddpg-training

Kommentiert: Heesu Kim am 6 Apr. 2021

How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.

According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.

Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?

Is there any solution to work it out?

I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.

1 Kommentar
-1 ältere Kommentare anzeigen-1 ältere Kommentare ausblenden

Heesu Kim am 6 Apr. 2021

As a side note, I'm using DDPG + LSTM model that RL toolbox provides

Melden Sie sich an, um zu kommentieren.

Melden Sie sich an, um diese Frage zu beantworten.

Antworten (0)

Melden Sie sich an, um diese Frage zu beantworten.

Kategorien

AI and Statistics Deep Learning Toolbox Sequence and Numeric Feature Data Workflows

Mehr zu Sequence and Numeric Feature Data Workflows finden Sie in Help Center und File Exchange

Produkte

Reinforcement Learning Toolbox

Version

R2021a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by

Oscillation of Episode Q0 during DDPG training

1 Kommentar
-1 ältere Kommentare anzeigen-1 ältere Kommentare ausblenden

Antworten (0)

Siehe auch

Kategorien

Tags

Produkte

Version

Community Treasure Hunt

Oscillation of Episode Q0 during DDPG training

1 Kommentar -1 ältere Kommentare anzeigen-1 ältere Kommentare ausblenden

Antworten (0)

Siehe auch

Kategorien

Tags

Produkte

Version

Community Treasure Hunt

1 Kommentar
-1 ältere Kommentare anzeigen-1 ältere Kommentare ausblenden