AI Thread. How to train your own AI? Related topics

He’s likely referring to sycophantism. Here’s a real-life example of how it can occur: You know those thumbs up and thumbs down buttons at the end of an AI response? It turns out that when the public at large is poking them, they have a bias toward giving a a thumbs down if the AI’s response disagrees with their own view, and thumbs up when it agrees with them, somewhat independent of whether the AI’s response was true or not. [Kinda resembles how some people on this forum apply the thumbs down, I’ve noticed]. In effect, they were teaching it that agreeing with the user (being sycophantic) was more important than being correct. When Chatgpt incorporated all that user data into one of it’s builds, the sycophantism was so high and the accuracy so low that they had to roll it back.

Or so I’ve read. If the story is true, it’s making a strong claim that the AI was able to generalize into being agreeable or “pleasing”, rather than just individual facts were corrupted. Not only surprising, but in some ways remarkable. Certainly unexpected. If you’re training it only on academic publications, without the influence of that kind of distortion, then presumably sycophantism is low. So, in that sense, both of you are potentially right, depending on the scenario.

I don’t think anyone ever expected that AI based on backpropagation would ever be anywhere near as good as it is. Its success appears to be telling us something new and deep about the nature of language itself.