The Frontier edition

The Frontier · 1 Oct 2026

1 Oct 2026 · 15 articles

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed…

Summary published by Simon Willison's Weblog · 29 Sep 2026 · Read the original ↗

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... They're not notably different from the GPT-6 family pelicans…

Summary published by Simon Willison's Weblog · 29 Sep 2026 · Read the original ↗

OpenAI DevDay 2026 live blog

I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the keynote. Tags: ai , openai , generative-ai , llms , coding-agents , live-blog , openai-devday

Summary published by Simon Willison's Weblog · 29 Sep 2026 · Read the original ↗

Claude Sonnet 5.5

New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of…

Summary published by Simon Willison's Weblog · 28 Sep 2026 · Read the original ↗

It's Time to Investigate the AI Labs

Article URL: https://calnewport.com/its-time-to-investigate-the-ai-labs/ Comments URL: https://news.ycombinator.com/item?id=49883471 Points: 620 # Comments: 276

Summary published by Hacker News - Newest: "AI" · 28 Sep 2026 · Read the original ↗

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything…

Summary published by Simon Willison's Weblog · 27 Sep 2026 · Read the original ↗

The Quest for Embedded Evaluators

Dario Amodei’s essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening.

Summary published by Don't Worry About the Vase · 27 Sep 2026 · Read the original ↗