Discussion about this post

User's avatar
Roy Xing's avatar

I have to wonder if RL would really be able to move beyond the look-up table interpretation though. Since RL still has the problem of exploration (or coverage). From what I've heard there was a fork in RL for robotics in which there was a question of if pure RL policies should be task-based or if they should be tracking-based. Task-based in the end lost out because it's much harder to solve the problem of exploring enough to learn a good behavior without getting stuck in a local optima. Now it seems like you either need human demonstrations to "warm-start" the policy or you go with RL tracking policies. Perhaps I'm misinterpreting your idea at the end though! I'd love to hear more

Tom Greenhaw's avatar

When you consider that LLMs use vectors to represent knowledge, apparent “grokking” or even novel creativity is expressed - emphasis on “apparent”. If the vectors are rotated, training data for one domain of knowledge can be applied to another topic unrelated to the original. It may appear novel and creative but really is based on the original data.

Additionally, if a body of knowledge on a given topic is considered an ingredient to a recipe, combinations with other ingredients can result in something unrecognizable from the original data. A cake does not resemble milk, flour, eggs and sugar and in the same way combinations of training data can appear novel.

1 more comment...

No posts

Ready for more?