Showing posts with label LLM. Show all posts
Showing posts with label LLM. Show all posts

8.05.2026

Qwen 3.8 Max IS OUT!

WorldofAI explores the capabilities of this large-scale model, focusing on its performance in coding, research, and long-horizon reasoning. The review examines its multimodal functionality and utility in generating 3D assets, web applications, and interactive simulations, while demonstrating the test-sprite CLI tool as a mechanism for verifying agent-driven code outputs.



7.31.2026

ThinkingCap - The Local Coding Model

In this video, Sam Witteveen looks at ThinkingCap, a fine-tune of the Qwen3.6-27B model. BottleCap's goal was intentionally conservative: do not try to make the model smarter or teach it new capabilities. They wanted to preserve knowledge, reasoning ability, answer quality, conversational style, instruction following, and safety behavior.

The only thing they wanted to change was the amount of computation spent reaching an answer. In other words: keeping the same model, and making sure it overthinks less.



7.27.2026

Anthropic launches Opus 5

On Friday, Anthropic launched its Opus 5 model, the newest version of its long-standing heavyweight model. While smaller than Fable 5, the model will be both cheaper and less restrictive than Fable, likely making it preferable in most use cases.

Opus 5 is launching only two months after Opus 4.8, which became available on May 28. Mythos 5, Fable 5, and Sonnet 5 all launched in June, leaving only the lightweight Haiku model still waiting for an upgrade to the 5 series.

7.24.2026

Poolside drops Laguna S 2.1, an open-weight coding model
that beats rivals 10x its size

Poolside, the San Francisco AI lab that has spent most of its three-year existence quietly selling coding models to governments and defense agencies, released its most capable model to date — and made an unusually aggressive bet that radical transparency, not raw scale, is how a smaller lab competes at the frontier.

The model, Laguna S 2.1, is a 118-billion-parameter Mixture-of-Experts (MoE) system that activates only 8 billion parameters per token, supports a context window of up to 1 million tokens, and — according to benchmarks published by the company — matches or beats open models several times its size on agentic coding tasks. The weights are available immediately on Hugging Face under the permissive OpenMDW-1.1 license.

7.22.2026

Google releases three new Gemini models — but no 3.5 Pro

Google DeepMind has released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Gemini 3.6 Flash is Google’s “workhorse model” that promises improved capabilities in coding, knowledge work, and multimodal performance while reducing token usage by up to 17%, making it cheaper than its predecessor 3.5 Flash.

7.17.2026

China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems

Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 — a 2.8-trillion-parameter model that the company says is now the largest open-source AI model in the world, and one that benchmarks show performs neck-and-neck with the most powerful proprietary systems from Anthropic and OpenAI.

The release, timed to land just ahead of the 2026 World Artificial Intelligence Conference in Shanghai, is a dramatic escalation in the global AI arms race and a watershed moment for the open-source AI movement. It also marks a remarkable comeback for a company whose market position had eroded significantly over the past 18 months following DeepSeek's meteoric rise.

7.16.2026

Thinking Machines amps up its bet against one-size-fits-all AI
with its first open model, Inkling

Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, released its first in-house AI model Wednesday morning, called Inkling. And unlike the flagship models from OpenAI, Anthropic, or Google, it’s open-weight, meaning outside developers and companies can download it and modify it directly.

Inkling is a mixture-of-experts system with 975 billion total parameters, though it only draws on a fraction of that — about 41 billion — for any given task, a common design that keeps very large models faster and cheaper to run. It was trained on 45 trillion tokens of text, image, audio, and video, and reasons natively across all four.

7.13.2026

GPT-5.6 IS HERE! BEST AI Model Ever? Beats Fable,
Faster, & Cheaper!

OpenAI has officially launched the GPT-5.6 family with Sol, Terra, and Luna, and after extensively testing every model, it's safe to say OpenAI is back in a big way.



7.10.2026

Meta enters the crowded AI coding battle with Muse Spark 1.1

Meta publicly launched a new version of Muse Spark on Thursday, a multimodal AI model designed for agentic coding that aims to compete with similar products offered by OpenAI and Anthropic.

Spark 1.1, the first version of which was announced in April, can engage in multistep reasoning and handle complex processes, manage digital workflows, and deploy new features in enterprise systems.

6.30.2026

Meituan open sources LongCat-2.0, the 1.6T, near-frontier agentic coding model that's been leading OpenRouter

Chinese delivery app company Meituan officially unveiled LongCat-2.0 on GitHub, Hugging Face, and its native platform, unmasking the model as the computational engine behind "Owl Alpha," the anonymous stealth model that has spent the last two months commanding global developer charts on OpenRouter. 

Developed to fundamentally disrupt closed-source enterprise dominance in autonomous software engineering, the 1.6-trillion-parameter Mixture-of-Experts (MoE) system brings a native 1-million-token context window to the public domain under a highly permissive, enterprise grade, commercially viable MIT license.

6.29.2026

Introducing Ornith 1.0 - Agentic Coding LLMs

Sam Witteveen explores this new family of self-scaffolding models designed to generate their own task-specific harnesses alongside solutions. By utilizing a two-stage reinforcement learning process, these models aim to optimize both the coding environment and agentic trajectories, offering a versatile approach for handling complex local coding tasks without requiring human-authored scaffolds.



6.26.2026

Liquid AI's smallest model yet LFM2.5-230M beats models 4X its size at data extraction, can run 'anywhere'

Liquid AI, founded by former MIT computer scientists, today released its smallest AI language model yet, LFM2.5-230M, and enterprises would do well to consider it for their uses in data extraction and local deployment on smartphones, laptops and robotics.

This is a 230-million-parameter foundation model explicitly designed for on-device agentic workflows, and as Liquid states in its release blog post, that small size makes it possible to run nearly "anywhere." According to Liquid, it also outperforms models more than 4X its size on selected benchmarks, specifically doing better at data extraction than the 800 million parameter count Alibaba Qwen3.5-0.8B (Instruct) and 1-billion parameter Google Gemma 3 1B.

6.24.2026

VibeThinker 3B - Taking on Giant Models

In this video, I look at VibeThinker 3b and how it is beating some models that are 300x its size on certain benchmarks by improving its reasoning and chain of thought to be better for specific use cases.  While the model is not for production it shows what could be done with these techniques.



6.18.2026

Kimi K2.7 Code: BEST Open Source Model? REALLY Cheap and
Beats
 
Opus 4.8 and GPT 5.5?

Kimi K2.7 Code might be one of the most impressive open-weight coding models released so far. In this video, we fully test Moonshot AI’s latest coding-focused model and see how it compares against Opus 4.8 Max, GPT-5.5, Fable 5, Qwen 3.7 Max, Grok 4.3, and other top coding models.



6.17.2026

Z.ai’s open-weights GLM-5.2 beats GPT-5.5 on multiple long-horizon coding benchmarks for 1/6th the cost

Today, Chinese AI startup Z.ai (formerly Zhipu AI) announced the immediate release of GLM-5.2, a 753-billion parameter open-weights large language model (LLM) engineered specifically to dominate "long-horizon" autonomous coding and engineering tasks. 

Available immediately on Hugging Face, the Z.ai API, and more than 20 third-party coding environments, the model boasts a highly stable 1-million-token context window alongside enterprise subscription tiers starting at just $12.60 per month.

6.12.2026

Google's DiffusionGemma generates 256 tokens in parallel and self-corrects as it goes

Google's DiffusionGemma, released this week, is an open source experimental model that applies diffusion to text generation at production scale. Built on the Gemma 4 backbone and released under the Apache 2.0 license, it is the first diffusion language model natively supported in the open source vLLM inference platform. It generates a 256-token block in parallel rather than sequentially, with every token position attending to every other. Google says DiffusionGemma generates text up to 4x faster than standard models on GPUs.

6.10.2026

Anthropic brings Mythos to the masses with Claude Fable 5,
its most powerful generally available model ever

Anthropic today launched two new AI models — Claude Fable 5 and Claude Mythos 5 — marking the company’s first broad release of the powerful “Mythos-class” AI capabilities it previously made available only to participating organizations in its restricted cybersecurity program, Project Glasswing, which it announced two months ago.

The company says Fable 5, which is the version most users and developers will get starting today, exceeds every Claude model it has previously made generally available — featuring stronger performance across software engineering, knowledge work, vision, scientific research and long-running tasks.

6.04.2026

Google's new open source Gemma 4 12B analyzes audio, video —
and runs entirely locally on a typical 16GB enterprise laptop

While many AI open source model providers are pursuing larger and more powerful models, Google is still giving attention to the smaller, more local side of the market. Today, the tech giant released Gemma 4 12B, an 11.95-billion-parameter open-weights model with permissive Apache 2.0 license optimized to execute locally on a standard enterprise laptop using just 16GB of VRAM or unified memory.

That means those enterprise users looking to keep working with AI while on a flight without WiFi, or trying to keep it offline for security reasons, can now do so far more easily and at far less cost (free to download and operate).

6.02.2026

MiniMax M3 IS INSANE! BEST Opensource AI Model!

In this video, I fully test MiniMax M3, the new open-weight frontier model from MiniMax that combines coding, agentic reasoning, multimodal understanding, and long-context capabilities into one model. M3 supports up to a 1 million token context window, is natively multimodal from day one, and delivers some seriously impressive benchmark results across SWE-Bench Pro, BrowseComp, SVG-Bench, KernelBench Hard, OSWorld Verified, and more.

What makes this release even more insane is the pricing. MiniMax M3 is not only competing with models like Opus 4.7 and GPT-5.5, but in several benchmarks it actually beats them while being dramatically cheaper. MiniMax is also offering huge token plans, aggressive API pricing, and open-weight access, making this one of the most accessible frontier-level models available right now.



5.29.2026

Anthropic's Claude Opus 4.8 is here with 3X cheaper fast mode and near-Mythos level alignment

Anthropic today released Claude Opus 4.8, an upgrade to its flagship model that ships at the same price as its predecessor, alongside a dramatically cheaper "fast mode" tier and a new feature that lets the model spawn hundreds of parallel subagents for codebase-scale work.