Full Description
It's another week of new model releases from the big AI labs OpenAI and Anthropic. This week, in keeping with its space theme, OpenAI introduces GPT-6 Astra, which boasts state-of-the-art computer use (an area ChatGPT has lagged in for some time), in addition to doing great on all the usual benchmarking tests and coding tasks. All charts are up and to the right, but the rollout to users will be gradual so no one actually has access to it yet. Ever wondered how watermarking of LLM generated content works? It's a lot more complicated than you might think, and Jack shares the rabbit hole he fell into when he wondered the same. There's a lot of computing and math, and basically, the only way to know for sure if something was generated by an AI model, is to check against that particular model. Which almost begs the question of if (or when) it really matters that much? And not to be outdone, Anthropic released Mythos 5.1 and Fable 5.1 this week too. Better at coding, knowledge work and long running tasks, of course, but also some interesting bits of how Claude is helping advance scientific research and how Anthropic is making agents less likely to go rogue during cybersecurity trials. And unlike OpenAI, Fable 5.1 is available for public use now. Lightning News is equally AI heavy this week with: a questionably-designed dual floor cleaning Roomba, a website called Felony Bench that tracks how many felonies AI agents have committed by company (OpenAI currently leads Anthropic by 1 felony), and after AI coding editor Cursor's acquisition by SpaceX, OpenAI has announced it will no longer allow access to its models. What a time to be alive!