Discussion about this post

User's avatar
Roy Xing's avatar
1hEdited

I have to wonder if RL would really be able to move beyond the look-up table interpretation though. Since RL still has the problem of exploration (or coverage). From what I've heard there was a fork in RL for robotics in which there was a question of if pure RL policies should be task-based or if they should be tracking-based. Task-based in the end lost out because it's much harder to solve the problem of exploring enough to learn a good behavior without getting stuck in a local optima. Now it seems like you either need human demonstrations to "warm-start" the policy or you go with RL tracking policies. Perhaps I'm misinterpreting your idea at the end though! I'd love to hear more

1 more comment...

No posts

Ready for more?