8.31.2026

Z.ai’s GLM-5.3 goes open weight

Earlier in August, Z.ai, the Chinese AI lab behind the viral ox-alpha model that turned out to be GLM-5.3-Flash, launched its flagship GLM-5.3 model. On Friday, the company made the model’s weights available on Hugging Face, which Nvidia may soon own. Several third-party inference services already host it and make it available on services like OpenRouter, which Stripe will soon own.

As The New Stack’s Amanda Caswell reported when the model originally launched, Z.ai’s focus on training GLM-5.3 was on post-training. The result isn’t simply a large jump in benchmark performance over its predecessor, GLM-5.2, but a model that is often ahead of other Chinese open-weight models and can keep pace with current models from the large U.S. frontier labs.

8.21.2026

This Is How Forward Deployed Engineering Is Actually Done

AI LABS explores the rising demand for forward-deployed engineers who bridge the gap between complex business processes and practical AI implementation. The session outlines a five-step framework for auditing existing workflows, identifying automation opportunities, and building systems that integrate seamlessly into a company's current operations without disrupting established productivity.



8.20.2026

Qwen3.8-27B - 200 Tokens per Second

In this video, I look at the long awaited Qwen3.8-27B model.  Both what it can do and how to serve it at the maximum tokens per second



8.19.2026

New v1.2 Skills!

Skills v1.2.0 is out with major improvements: new documentation site, Claude Code marketplace integration, five brand new skills including Wait What for Opus verbosity, an updated Grill Me with multi-question rounds, and the powerful Wizard skill for infrastructure provisioning.



8.18.2026

Qwen3.8-27B runs frontier-class coding agents & reasoning locally

The biggest AI model release of the past few days, at least among the developers and AI power users on social media, wasn't a frontier cloud model from OpenAI, Anthropic or Google.

It was a 27-billion-parameter model from Alibaba: Qwen3.8-27B landed on Hugging Face on Friday under an enterprise-friendly, open source Apache 2.0 license, giving developers downloadable weights for a dense multimodal model.

8.17.2026

This Claude Skill Just Fixed Loop Engineering

AI LABS explores the mechanism behind the popular Gauntlet Loop, a technique for building complex applications in one shot. By integrating Wayfinder, an intensive planning skill, AI LABS demonstrates how to overcome the method's inherent limitations regarding quality verification and project drift, providing a more structured approach to AI-driven software development.



8.14.2026

Local AI On Apple Silicon uses 7X Less RAM

Better Stack explores how Turbo Fieldfare utilizes the unique architecture of Apple Silicon to run massive mixture-of-experts models. By streaming specific model components directly from SSD and leveraging unified memory, this approach significantly reduces RAM requirements, making advanced AI performance more accessible on local hardware.