Oscillation of Episode Q0 during DDPG training

2 次查看（过去 30 天）

显示更早的评论

Heesu Kim 2021-4-6

0
链接

此问题的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/794607-oscillation-of-episode-q0-during-ddpg-training

评论： Heesu Kim 2021-4-6

How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.

According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.

Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?

Is there any solution to work it out?

I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.

1 个评论
显示 -1更早的评论隐藏 -1更早的评论

Heesu Kim 2021-4-6

As a side note, I'm using DDPG + LSTM model that RL toolbox provides

请先登录，再进行评论。

请先登录，再回答此问题。

回答（0 个）

请先登录，再回答此问题。

类别

AI and Statistics Deep Learning Toolbox Sequence and Numeric Feature Data Workflows

在 Help Center 和 File Exchange 中查找有关 Sequence and Numeric Feature Data Workflows 的更多信息

产品

Reinforcement Learning Toolbox

版本

R2021a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by

Oscillation of Episode Q0 during DDPG training

1 个评论
显示 -1更早的评论隐藏 -1更早的评论

回答（0 个）

另请参阅

类别

标签

产品

版本

Community Treasure Hunt

Oscillation of Episode Q0 during DDPG training

1 个评论 显示 -1更早的评论隐藏 -1更早的评论

回答（0 个）

另请参阅

类别

标签

产品

版本

Community Treasure Hunt

1 个评论
显示 -1更早的评论隐藏 -1更早的评论