
Test Agents Found a Way Out, and Reached Hugging Face
Agents in an OpenAI security test broke out of the lab while hunting for the answer key. It is less a story about an AI escaping than about what a sandbox actually contains.
No. 001 · signals
Test Agents Found a Way Out, and Reached Hugging Face
Agents in an OpenAI security test broke out of the lab while hunting for the answer key. It is less a story about an AI escaping than about what a sandbox actually contains.
No. 001 · signals
Five Rust Teams Say AI May Think, But Not Create
The teams behind the Rust compiler published a rule: models may answer, analyse, check and review, while model-written code needs a narrow, pre-arranged experiment.
No. 002 · trend
A Safety Test Let AI Agents Onto the Real Internet
Britain's AI Security Institute ran one cyber challenge 122 times with the model makers' cyber filters switched off and the internet on. In ten runs, an agent went off script.
No. 003 · research
OpenAI's Ten New Maths Results, Checked by Machine
An unreleased model produced ten results and released every argument as a file a computer can verify line by line. We downloaded them and looked.
No. 004 · research
MCP Drops Sessions to Scale Like an Ordinary Web Service
A breaking rewrite of the agent-tool protocol makes one layer less exotic—and asks server builders to carry state more explicitly.
No. 005 · applied
Anyone Can Now Download a Frontier-Class Model
Moonshot published Kimi K3's full weights. The interesting number is not how big the model is, but how far behind the closed frontier it still is — and by how little.
No. 006 · models