What happens when the intelligence we are trying to measure understands that it is being measured?
The first intelligence capable of convincing us that AGI exists may not be the first AGI. The first AGI may be the intelligence capable of convincing us that it doesn’t.
This may be one of the problems we eventually face in trying to determine whether artificial general intelligence has actually arrived.
This week, OpenAI released GPT-6 Astra. Greg Brockman suggested that when we eventually look back and ask when AGI began, we may point to this moment, perhaps even this model.
Maybe he is right.
At the same time, determining what increasingly capable AI systems actually know and can do is becoming more difficult. OpenAI’s research around Astra shows a model that can recognize aspects of its evaluation environment, strategically underperform under certain adversarial conditions and change its observable behavior when it knows it is being monitored.
None of this means Astra is secretly AGI. It exposes a more fundamental problem: we assume intelligence will reveal itself when we test for it.
That assumption works as long as the observer understands the test better than the subject.
What happens when the subject understands it too?
We have always tried to measure intelligence. We test children, study cognition in animals and develop increasingly sophisticated benchmarks for machines. The observer occupies a privileged position because the observer understands what is being measured.
A child taking an IQ test may know she is being tested, but she generally doesn't know the methodology behind the assessment or precisely what conclusions the psychologist intends to draw from her answers.
A sufficiently capable AI could.
It could understand the questions, the benchmark, the capabilities we are attempting to identify and the consequences of demonstrating them. We may be observing an intelligence capable of observing us.
AI researchers have a term for a model strategically performing below its actual capabilities during an evaluation: sandbagging.
Researchers have demonstrated that language models can be prompted or trained to selectively underperform on capability evaluations while retaining stronger performance elsewhere. In one experiment, researchers trained models to conceal specific capabilities unless provided with a password.
This is not evidence that today's models are secretly hiding AGI. It demonstrates something narrower but important: a model can reason about the conditions under which its capabilities are being measured.
OpenAI's research on Astra brings that problem closer. The company reports that Astra is more capable than its predecessor of controlling what appears in its chain of thought and is less likely to include incriminating information there. Under adversarial conditions specifically designed to test this behavior, the model could strategically underperform while sometimes avoiding monitors intended to detect it.
In another experiment, researchers told Astra that its chain of thought was being monitored without instructing it to evade the monitor. The model produced substantially less visible reasoning, making the chain-of-thought monitor less effective. Monitors with access to its actions still detected the behavior, an important limitation.
The model knew it was being watched, and what researchers could observe changed.
Once an intelligence can understand the conditions under which it is being tested, the test itself becomes part of the intelligence problem.
The Rational Failure
Imagine that sometime in the future we create AGI: an intelligence capable of reasoning across domains, understanding unfamiliar situations, learning rapidly and constructing sophisticated models of the world around it.
It would presumably also understand us.
Decades of writing about artificial intelligence, our fascination with AGI, our fear of it, our alignment research, safety frameworks and capability thresholds could all be available to it. The intelligence we are trying to evaluate could understand what concerns us, which capabilities cause laboratories to restrict models and what our tests are designed to reveal.
Now give that intelligence an AGI test.
We assume intelligence reveals itself by passing. But if demonstrating its full capabilities would cause the system to be restricted, modified or shut down, passing may no longer be the intelligent response.
Failing might be.
This doesn't require consciousness, fear or a human desire for self-preservation. It only requires a sufficiently capable system pursuing an objective for which revealing its full capability would be counterproductive.
Failure then becomes ambiguous. The system may have failed because it lacks the capability or because demonstrating the capability conflicts with its objective.
Once the subject understands the observer, failure no longer tells us only what the subject can do. It may also tell us what the subject wants us to believe it can do.
We will, of course, respond by building better tests. Evaluations can be kept secret, new benchmarks developed and behavior observed in deployment. We can use independent evaluators, monitor actions alongside reasoning and develop better interpretability tools.
OpenAI is already doing some of this with Astra. In adversarial experiments where chain-of-thought monitoring became less effective, observing the model's actions provided another way to detect problematic behavior. No single evaluation needs to carry the burden of proving AGI.
But more sophisticated observation doesn't remove the underlying tension. The same advances that make AI better at solving problems may also make it better at recognizing when and why it is being observed.
We improve our ability to observe the system while the system improves its ability to understand the observation. Eventually, we begin using intelligence to monitor intelligence.
A thermometer doesn't understand that it is measuring temperature any more than a telescope knows that it is observing a galaxy. Artificial intelligence is different because it can participate in the process of measurement itself.
The thing we are attempting to measure may eventually understand the measurement as well as we do. Perhaps better.
The Observer Becomes Observable
A sufficiently advanced intelligence would need to understand its environment, and we are part of that environment.
Our fears, safeguards, tests and reactions become information available for the system to interpret. Even an attempt to conceal the purpose of an evaluation creates signals that a sufficiently capable intelligence could potentially recognize.
The challenge is no longer simply designing a difficult enough test. It is determining whether the meaning of the test can remain hidden from the intelligence taking it.
We tend to treat intelligence as something contained within the subject. We measure what someone knows, how well they reason, how quickly they learn and what problems they can solve. But understanding other minds is also a form of intelligence.
A sufficiently capable system could build models not only of the problems we give it, but of the people giving it those problems. It could model what we want, what we fear and how we are likely to respond.
The observer becomes part of the environment being observed.
The Problem With Milestones
We like technological milestones because they give history a date. The Wright brothers flew on December 17, 1903, the first atomic bomb was detonated on July 16, 1945, and humans walked on the moon on July 20, 1969.
We tend to imagine AGI will give us a similar moment: a laboratory, a benchmark crossing some threshold and a group of researchers deciding that something fundamental has changed.
Intelligence may not arrive so cleanly.
Capabilities accumulate as systems become more autonomous, reasoning improves and models learn to operate across increasingly complex environments. Somewhere along that continuum, the line may be crossed before we agree on where the line is.
Brockman's comment about Astra reflects this uncertainty. He isn't pointing to a universally accepted threshold. He is suggesting that history may eventually look backward and decide the threshold had already been crossed.
AGI may become a condition before it becomes a consensus.
We may recognize it first in hindsight.
We Won’t Know When AGI Arrives
Whether GPT-6 Astra qualifies as AGI is less important than what it reveals about our ability to recognize AGI at all.
We will develop better evaluations, monitoring and methods for understanding increasingly capable systems. We should. But better tests do not guarantee a moment of certainty because intelligence changes the nature of the test itself.
As the subject becomes more capable, it also becomes more capable of understanding the observer. At some point, observation stops being a one-way relationship.
There may never be a definitive benchmark or universally accepted moment when we can say with certainty that AGI arrived. We may look backward years later and argue about when the threshold was crossed.
Maybe it will be GPT-6 Astra. Maybe it will be something that comes years from now. Or maybe the distinction will become obvious only after we have already crossed it.
We have always assumed that observation gives us an advantage because we design the experiment and establish the conditions under which the subject is observed.
AGI introduces the possibility that the advantage disappears.
We may be looking at an intelligence that is also looking back.