
Evaluating AI Agents: Accuracy Is the Wrong Metric
Agents have no single correct path, succeed partially, and behave differently run to run. What to measure instead of accuracy, and why failure shape matters most.

Agents have no single correct path, succeed partially, and behave differently run to run. What to measure instead of accuracy, and why failure shape matters most.

You cannot prevent an agent from making a bad decision. You can make sure a bad decision cannot do much. A practical security model built around tools, not prompts.

OpenAI’s published tables put Terra and Luna within a tenth of a point on professional agentic work and 48 points apart on long-context recall. Route by failure mode, not difficulty.

SQL injection was solved by separating code from data. Language models cannot make that separation, which is why prompt injection defences have to be architectural.

The useful distinction is not how many model calls a system makes, but whether it decides its own control flow. That difference changes testing, failure modes and security.

An AI model halved HAWK’s effective key strength in 60 hours. But the finding that runs end to end on a desktop, against a deployed ISO/IEC cipher, is the one worth reading twice.

Jensen Huang’s first X post shared the Open Weights and American AI Leadership letter. What open weights mean, why Washington is arguing, and who gets priced out.