Blog
Latest posts, updates, and notes.

Jev: The Promise and Limits of a Model Built to Decide
TypeSafe built Jev to make decisions, not write text. We explore how it works and test the company's claims against inde…
Read more
Spend on the harness before the judge
Anthropic's red-team study of Auto Mode shows how an LLM judge's harness shapes what the judge sees.
Read more
Anthropic's Hacker-Opus: When the Checker Becomes the Task
Anthropic's Hacker-Opus experiment shows how visible graders, score-focused prompts, and agent loops can push reward-see…
Read more