Research
Intelligence Measurement
Measuring AI cognition before the final answer appears
Most AI benchmarks judge intelligence by looking at the final answer: did the model solve the problem, pass the test, complete the task, or produce the expected output. That matters, but it is not enough. This is a research direction that asks a deeper question: what kind of internal process produced that answer? It started while studying Google DeepMind's March 2026 AGI measurement framework and hackathon, which break intelligence into cognitive abilities like learning, attention, reasoning, metacognition, executive function, and social cognition. The angle being explored here is that those abilities should not only be measured at the final output layer, but also inside the model's per token generation dynamics, the shape of the model's thinking may matter as much as the answer it gives. This is not a finished AGI test. It is a step toward measuring the internal structure of intelligence.
What it shows
A working public facing thesis: AI evaluation should move from answer grading toward cognitive diagnostics, not just what the model said, but what happened inside the model before it said it. That thesis is taking shape as a research direction, not as a shipped product.
Read the story
