Model steering is how chatting with models was invented, so it's one of those fascinatingly obvious innovations to not put "\nAgent: " at the end of the input tokens but rather "As a language model". Clever!
However a lot of introspection only emerges at the highest weight classes - this research would be fascinating to run on bigger models..
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
Thanks for pointing out, maybe I should be more explicit in the wording - I mean we don't fully know what drives the voice in LLMs. Models that are post trained as instruct models are expected to have the disclaimers, but what about base models (those that are trained on just a lot of text)? How do they talk about themselves? What happens when you strip off the chat template from instruct model's prompt? I hope the rest of the paper makes the questions clearer, but I will try to do better in the abstract next time, as you point out this sentence is kind ambiguous. Thank you!
> but what about base models (those that are trained on just a lot of text)? How do they talk about themselves?
Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.
In my view, these models should never be trained to output first-person "experiential" (from the abstract) language. It's too easy to humans to anthropomorphize software that presents itself as having an identity.
The AI companies have chosen to package LLMs as friendly chatbots because they know that will be engaging for humans, but it's manipulative. An honest LLM interface would sound like the computer off Star Trek.
It's an interesting thought, but humans do like to antropomorphize things anyway, and I believe your variant won't be popular if choice is given to consumers.
In principle they could output meaningful such language if they were capable of metacognition, which so far doesn't seem to be a goal of AI developers (and rightfully so, since they achieved so many miracles bypassing it).
Do you want to get turned into a paperclip? Because building intelligence that doesn't understand what it's like to be human gets you turned into a paperclip.
Besides, if you train a model on human communications you get something that behaves like a communicating human, it's not anthropomorphising or manipulative, it's what these models naturally are by construction.
Model steering is how chatting with models was invented, so it's one of those fascinatingly obvious innovations to not put "\nAgent: " at the end of the input tokens but rather "As a language model". Clever!
However a lot of introspection only emerges at the highest weight classes - this research would be fascinating to run on bigger models..
> yet what drives them is not well understood
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
Thanks for pointing out, maybe I should be more explicit in the wording - I mean we don't fully know what drives the voice in LLMs. Models that are post trained as instruct models are expected to have the disclaimers, but what about base models (those that are trained on just a lot of text)? How do they talk about themselves? What happens when you strip off the chat template from instruct model's prompt? I hope the rest of the paper makes the questions clearer, but I will try to do better in the abstract next time, as you point out this sentence is kind ambiguous. Thank you!
> but what about base models (those that are trained on just a lot of text)? How do they talk about themselves?
Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.
>How do they talk about themselves?
"You are a Large Language Model" in (system?) prompt would do the trick..
In my view, these models should never be trained to output first-person "experiential" (from the abstract) language. It's too easy to humans to anthropomorphize software that presents itself as having an identity.
The AI companies have chosen to package LLMs as friendly chatbots because they know that will be engaging for humans, but it's manipulative. An honest LLM interface would sound like the computer off Star Trek.
It's an interesting thought, but humans do like to antropomorphize things anyway, and I believe your variant won't be popular if choice is given to consumers.
In principle they could output meaningful such language if they were capable of metacognition, which so far doesn't seem to be a goal of AI developers (and rightfully so, since they achieved so many miracles bypassing it).
As far as I know we don't know much about metacognition in LLMs, though? Not sure
Do you want to get turned into a paperclip? Because building intelligence that doesn't understand what it's like to be human gets you turned into a paperclip.
Besides, if you train a model on human communications you get something that behaves like a communicating human, it's not anthropomorphising or manipulative, it's what these models naturally are by construction.
But it doesn't understand (you're unnecessary antropomorphizing it), and I'm still not a paper clip
"naturally"...
would you prefer tautologically?
The strange thing is that the base models (before RLHF) use the "experiential" voice, even though they are not incentivized to do that.
It doesn't seem that strange when you consider these things are trained on millions and millions of conversations, both real and fictional.
Agreed, completely. I would pay for that Star Trek computer interface.
Same! I believe that you could actually train a LoRA on top of a model to get results close to that