LangChain announced new updates to LangSmith. Updates include Engine v2 with red teaming and automatic testing, a new version of Managed Deep Agents, trajectories and more.
Summary published by LangChain Blog · 25 Sep 2026 · Read the original ↗
Google DeepMind introduces SynthID Bio to watermark AI-designed proteins while maintaining biological function. Read the full research report here.
Summary published by Google DeepMind · 30 Sep 2026 · Read the original ↗
Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests.
Summary published by LangChain Blog · 29 Sep 2026 · Read the original ↗
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed…
Summary published by Simon Willison's Weblog · 29 Sep 2026 · Read the original ↗
My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... They're not notably different from the GPT-6 family pelicans…
Summary published by Simon Willison's Weblog · 29 Sep 2026 · Read the original ↗
I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the keynote. Tags: ai , openai , generative-ai , llms , coding-agents , live-blog , openai-devday
Summary published by Simon Willison's Weblog · 29 Sep 2026 · Read the original ↗
New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of…
Summary published by Simon Willison's Weblog · 28 Sep 2026 · Read the original ↗
Article URL: https://calnewport.com/its-time-to-investigate-the-ai-labs/ Comments URL: https://news.ycombinator.com/item?id=49883471 Points: 620 # Comments: 276
Summary published by Hacker News - Newest: "AI" · 28 Sep 2026 · Read the original ↗
Article URL: https://www.cnbc.com/2026/09/28/nvidia-releases.html Comments URL: https://news.ycombinator.com/item?id=49879883 Points: 227 # Comments: 299
Summary published by Hacker News - Newest: "AI" · 28 Sep 2026 · Read the original ↗
Article URL: https://thecivilian.co.nz/2026/09/27/ai-companies-in-fierce-arms-race-to-demonstrate-their-model-is-the-most-existentially-threatening-to-humanity/ Comments URL: https://news.ycombinator.com/item?id=49875148 Points: 440 # Comments: 395
Summary published by Hacker News - Newest: "AI" · 28 Sep 2026 · Read the original ↗
Article URL: https://arxiv.org/abs/2110.01834 Comments URL: https://news.ycombinator.com/item?id=49873241 Points: 177 # Comments: 82
Summary published by Hacker News - Newest: "AI" · 28 Sep 2026 · Read the original ↗
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything…
Summary published by Simon Willison's Weblog · 27 Sep 2026 · Read the original ↗
Dario Amodei’s essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening.
Summary published by Don't Worry About the Vase · 27 Sep 2026 · Read the original ↗
See how LangGraph orchestrates Jev, TypeSafe AI's decision model, to build faster, cheaper production agents.
Summary published by LangChain Blog · 25 Sep 2026 · Read the original ↗
What is Jev? Learn how TypeSafe AI’s System One model makes fast, structured decisions, where it fits in the agent loop, and how to use Jev with LangChain
Summary published by LangChain Blog · 20 Sep 2026 · Read the original ↗