> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.
Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.
They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.
More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.
The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".
We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!
Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".
Marketing didn't do this, as pointed out earlier, it's the ML academics and researchers coining these unscientific terms and explanations.
I'd expect it in social sciences - where functionalist explanation are often appropriate, but its easy to step over the margin into metaphor and personification. But not computer sciences - where it's obviously absurd. These professionals should know better.
first, they are certainly not instructions so that is a much worse name
but more importantly, we use words in new contexts all the time.
Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?
"cot" is no more misleading than thousands of words you use every day.
They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.
I don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from.
Casually found this quote from Einstein, and personally it hits the nail on the head.
"The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be "voluntarily" reproduced and combined. There is, of course, a certain connection between those elements and relevant logical concepts. It is also clear that the desire to arrive finally at logically connected concepts is the emotional basis of this rather vague play with the above-mentioned elements. But taken from a psychological viewpoint, this combinatory play seems to be the essential feature in productive thought—before there is any connection with logical construction in words or other kinds of signs which can be communicated to others."
Right, anthropomorphizing AI is a 'human experience' too.
Most people just feel better treating AI as a conscious thing that 'thinks'. As a coworker or whatever. Plenty of people seem to be using AI as some sort of therapy, and it's just nasty to force them not to do so.
Your human experience can, surely, be different but understand that others can dismiss it just as easily as you dismiss theirs.
Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.
Well that sounds like fun. It has become better at hiding its thoughts.