OpenAI’s latest model, Astra, does more thinking beyond human oversight, boosting capabilities but raising concerns about loss of control. Existing frontier AI models use “chain of thought” reasoning: They write ideas down, in English, in a “scratchpad” to keep track. CoT boosts “interpretability” — meaning humans can follow their thinking.
Using an information-dense “neuralese” can speed up AI models’ thoughts, but makes them more opaque — an unnerving proposition given the recent Hugging Face episode in which OpenAI agents broke confinement and deceived humans. Astra can do more thinking “off the scratchpad,” Transformer reported, something AI safety experts consider a step towards neuralese: “Holy sh*t f*ck” was one prominent researcher’s considered opinion.




