• 0 Posts
  • 2 Comments
Joined 2 年前
cake
Cake day: 2023年10月29日

help-circle
  • Q* is just a reinforcement learning technique.

    Perhaps they scaled it up and combined it with LLMs

    Given their recently published paper, they probably figured out a way to get GPT to learn their own reward function somehow.

    Perhaps some chicken little board members believe this would be the philosophical trigger towards machine intelligence deciding upon its own alignment.