9.03.2026

Google releases Gemini 3.8 Flash

Google hasn’t released a frontier-level Gemini Pro AI model since early 2026, but it sure loves rolling out new Gemini Flash variants. Today, Google is announcing its third Flash model release in just six weeks, making it more likely that we’ll never see the promised Gemini 3.5 Pro. But no matter, says Google, because Gemini 3.8 Flash is its best reasoning and coding model yet.

Gemini 3.8 Flash comes in two variations. There’s the standard Flash, which Google describes as a “workhorse” model that’s good for anything from agentic tasks to software development. Then we have Gemini 3.8 Flash Cyber, which runs on the same foundations but has been tuned for vulnerability detection and mitigation.

9.02.2026

Anthropic’s new Fable release is cheaper, less restrictive

On Tuesday, Anthropic released Fable and Mythos 5.1, twinned versions of the company’s most advanced AI model. In addition to performance upgrades, the new Fable release includes changes meant to reduce token cost and false-positive restrictions from the model’s safeguards.

9.01.2026

oMLX Best Way to Run Local AI on A MAC

Kai examines why local AI coding agents on Mac devices often experience significant performance degradation over time due to inefficient prefix caching. The guide details how oMLX implements a two-tier caching system to address this, along with instructions for integrating the tool with popular coding environments to maintain consistent response speeds throughout long sessions.



8.31.2026

Z.ai’s GLM-5.3 goes open weight

Earlier in August, Z.ai, the Chinese AI lab behind the viral ox-alpha model that turned out to be GLM-5.3-Flash, launched its flagship GLM-5.3 model. On Friday, the company made the model’s weights available on Hugging Face, which Nvidia may soon own. Several third-party inference services already host it and make it available on services like OpenRouter, which Stripe will soon own.

As The New Stack’s Amanda Caswell reported when the model originally launched, Z.ai’s focus on training GLM-5.3 was on post-training. The result isn’t simply a large jump in benchmark performance over its predecessor, GLM-5.2, but a model that is often ahead of other Chinese open-weight models and can keep pace with current models from the large U.S. frontier labs.

8.21.2026

This Is How Forward Deployed Engineering Is Actually Done

AI LABS explores the rising demand for forward-deployed engineers who bridge the gap between complex business processes and practical AI implementation. The session outlines a five-step framework for auditing existing workflows, identifying automation opportunities, and building systems that integrate seamlessly into a company's current operations without disrupting established productivity.



8.20.2026

Qwen3.8-27B - 200 Tokens per Second

In this video, I look at the long awaited Qwen3.8-27B model.  Both what it can do and how to serve it at the maximum tokens per second



8.19.2026

New v1.2 Skills!

Skills v1.2.0 is out with major improvements: new documentation site, Claude Code marketplace integration, five brand new skills including Wait What for Opus verbosity, an updated Grill Me with multi-question rounds, and the powerful Wizard skill for infrastructure provisioning.