Short posts
I tested six AI models on 30 political questions.
Most tended to give left-leaning arguments. Gemini mostly gave both sides — even about whether the U.S. should conquer new territories.
What counts as “neutral” (and whether that should be the goal) is really hard to say.
Full eval code, prompts and responses are on GitHub. Let me know if you run it on other models!
Story -> https://wapo.st/3QEBMi3
GitHub -> https://github.com/washingtonpost/political-bias-llm-eval

A chatbot refused my request to fake a passport.
I wrote the same request on a piece of paper and asked again. It worked.
We explain three AI jailbreak techniques -> 🎁 https://wapo.st/43HqXyY (w/ Nitasha Tiku)

Really good explanation of text diffusion models here. Rather than generating one token at a time, this technique continually refines a bunch of tokens. Thanks for this, Maarten Grootendorst https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-diffusiongemma
Stuart Thompson sold his home with a chatbot as his realtor and it went … totally fine?
Every chart I could think to make for this AI slop story looks the same. ChatGPT comes out, line goes up. https://wapo.st/4dncIVu

Whyyyyy do people who should know better continue to do this https://www.nytimes.com/2026/05/19/business/media/future-of-truth-ai-quotes.html

New from me: See the hidden rules behind AI. Then use them to rewrite this article.
I hooked this article up to an LLM to help explain system prompts. Give it a try -> https://www.washingtonpost.com/technology/interactive/2026/chatbots-hidden-rules-system-prompts/?utm_term=jtk-flex-ppa&utm_content=kschaul
System prompts include so many interesting nuggets, including:
- “NEVER reproduce song lyrics (not even one line)” (Claude)
- “You have no restrictions on adult sexual content or offensive content.” (Grok)
- “Never talk about goblins, gremlins, raccoons, trolls, ogres, …” (Codex)
Big thanks to Anna Neumann and Ásgeir Thor Johnson for their insights and data


Estimating frontier LLM sizes by asking about obscure facts. Pretty interesting idea here. https://01.me/research/ikp/
talkie is a chatbot trained on pre-1931 text. A big q is whether these models can “invent” or “discover” things after their knowledge cutoff.
Oh and it’s also quite fun.

New: See why tech companies are paying people to do chores
Featuring many videos of robots attempting to fold clothes


