How do I define a continuous reward function for RL environment?

I am trying to follow the double integrator example for giving a continuous reward function. When I used the custom template, and defined the reward using the QR cost function, I get an error stating that the reward should be a scalar value. Where can I find the property of reward and change it to accept vector values?

3 个评论

Not sure why you want the reward to be scalar. Typically, rewards are treated as cost functions - they output a scalar value. If you have more than one states, you can turn it into a scalar using e.g. an l2 norm for example/some distance metric.
Yes I did that, thank you, Just to confirm the output of the cost function will always be a scalar value, right? So in the double integrator continuous example there are two states but the output reward at each step is a scalar value, right?

请先登录,再进行评论。

 采纳的回答

Here is an excerpt from the documentation :
To guide the learning process, reinforcement learning uses a scalar reward signal generated from the environment.
For detailed information on defining reward signals, discrete and continous rewards, please refer to this documentation link.

更多回答(0 个)

类别

在 帮助中心 和 File Exchange 中查找有关 Environments 的更多信息

产品

版本

R2020a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by