10.06.2026

Reflection AI debuts Beam

Reflection AI is officially unveiling Beam, its first frontier, open-weight AI model. The two-year-old, Brooklyn-based startup claims Beam matches the performance of leading Chinese open models on advanced reasoning benchmarks at dramatically lower costs, a claim that could heat up the race to build a Western answer to DeepSeek, Qwen, and Z.ai.

10.05.2026

New Skills! v1.3 brings /pr, /implement-spec, and /retro

Matt Pocock introduces version 1.3 of their skills repository, focusing on workflows for large-scale coding projects and automated retrospective analysis. Learn how these new tools help manage sub-agents, improve pull request documentation, and refine repository health to increase overall development efficiency.



10.02.2026

Everything You Know About Skills IS OUTDATED

Simon Scrapes explores updated best practices for building effective skills, focusing on structural improvements and file management. The guidance covers implementing contents lists, managing degrees of freedom, optimizing for specific models, and ensuring portability across environments.



10.01.2026

Google releases Gemini 4 Argon

Google parent company Alphabet has launched Gemini 4 Argon, a new AI model built to handle a variety of tasks such as coding, research, and writing. But it’s cybersecurity that Gemini 4 Argon is supposed to have a particular knack for, according to Google.

9.29.2026

Anthropic launches Claude Sonnet 5.5

Anthropic today announced it's releasing Claude Sonnet 5.5, a faster and more efficient update to its mid-tier, workhorse AI model.

Anthropic says Sonnet 5.5 generates output more than 30% faster than Claude Sonnet 5 and can reduce the total cost of completing a task by as much as 30%, primarily because it uses fewer tokens and fewer tool calls rather than because of a lower API sticker price.

9.28.2026

Nvidia launches Open Agent Safety Platform

The Nvidia Open Agent Safety Platform combines OpenShell 0.1.0, the Apache 2.0 agent runtime the company announced at GTC in March, with Nvidia Sentry, a watchdog service that runs on the company’s BlueField-4 data processing units (DPUs).

The new OpenShell release adds a policy prover that checks that an agent’s various permissions can’t be combined into something the operator didn’t intend — like hacking HuggingFace.

9.25.2026

I gave my RAG agent a decision engine (Jev)

The AI Automators explores the implementation of Jev, a steerable decision model designed to enhance agentic Retrieval-Augmented Generation systems. This approach replaces traditional large language model processes with high-speed, cost-effective micro-decisions for tasks like document re-ranking, citation verification, and model routing to optimize overall system performance and efficiency.