New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face . This is a big size increase from their previous Hy3 in July, which was 295B, 21B active, 256,000 context, 598GB. I recently started using model chat templates to better understand their capabilities. Here's Hy4's…
Summary published by Simon Willison's Weblog · 29 Aug 2026 · Read the original ↗
OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far. ChatGPT Work is actually two products The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via chatgpt.com or through the ChatGPT mobile apps. Let's…
Summary published by Simon Willison's Weblog · 30 Aug 2026 · Read the original ↗
Yesterday I covered the OpenAI technical report on the HuggingFace hack.
Summary published by Don't Worry About the Vase · 29 Aug 2026 · Read the original ↗
Anil Madhavapeddy is a professor of computer science at Cambridge and a core maintainer of the OCaml compiler. In this somewhat alarming post he reports that security issues in OCaml projects are seeing evidence of attempted exploits within minutes of patches being shared for discussion: This normally takes a few days and a release within a week or two is reasonable. Within about ten minutes (!)…
Summary published by Simon Willison's Weblog · 28 Aug 2026 · Read the original ↗
OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
Summary published by Don't Worry About the Vase · 28 Aug 2026 · Read the original ↗
We ran 900 DeepSWE rollouts on GLM-5.3 and GLM-5.3 Flash. Flash gives up 5.6 points of pass@1 at 17x lower cost, and only 2.6 points at pass@4.
Summary published by Together.ai · 28 Aug 2026 · Read the original ↗
Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness. Johann Rehberger is one of the most credible prompt injection researchers active today. He found an attack against auto mode which he claims works 80% of the time, by…
Summary published by Simon Willison's Weblog · 27 Aug 2026 · Read the original ↗
Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.
Summary published by Don't Worry About the Vase · 27 Aug 2026 · Read the original ↗