Short posts
This is a must-read. Such important reporting, told so well. 🎁 https://wapo.st/3RoU0Vu
How long did OpenAI’s hacking agent run rogue before the company noticed? We mapped out what we know.
Beware: You’ll have to scroll a while.
Dig in, folks
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Two thoughts on the OpenAI-HuggingFace hack:
- If OpenAI cannot secure their models, maybe they should not be building them
- This hacking technology exists in AI labs, and to a growing extent outside them too. This is now our reality.
OpenAI says its model hacked out of its sandbox, then hacked HuggingFace to find answers to an eval: OpenAI and Hugging Face partner to address security incident during model evaluation
HuggingFace previously wrote about it here: Security incident disclosure — July 2026
Why did OpenAI not sufficiently secure its training environment? Weird humble-brag vibe going on. I hope we get more details on the exploits soon.
The AI 2027 folks are back, this time arguing for an international deal to slow down AI development https://ai-2040.com/
Anthropic says its weakest model, many of OpenAI’s and at least one Chinese open-weight model can all find the vulns and exploit that prompted the U.S. to effectively ban it. https://www.anthropic.com/news/redeploying-fable-5
Pretty crazy how quickly the Trump admin reversed itself on AI https://wapo.st/4xS9cui
I tested six AI models on 30 political questions.
Most tended to give left-leaning arguments. Gemini mostly gave both sides — even about whether the U.S. should conquer new territories.
What counts as “neutral” (and whether that should be the goal) is really hard to say.
Full eval code, prompts and responses are on GitHub. Let me know if you run it on other models!
Story -> https://wapo.st/3QEBMi3
GitHub -> https://github.com/washingtonpost/political-bias-llm-eval
A chatbot refused my request to fake a passport.
I wrote the same request on a piece of paper and asked again. It worked.
We explain three AI jailbreak techniques -> 🎁 https://wapo.st/43HqXyY (w/ Nitasha Tiku)