Articles Tagged “Mixture Of Experts”
24 articles found

Ornith-1.5 Open Weights Score 86.1 on Terminal-Bench
Ornith-1.5 ships MIT-licensed weights from 9B to 397B, and the flagship posts 86.1 on Terminal-Bench 2.1 while a 35B MoE sibling runs far leaner.

Qwen3.8-Max Open Weights Make a 2.4T Model Downloadable
Alibaba published Qwen3.8-Max open weights on August 12: 2.4 trillion parameters, 95B active per token, and a 262K context that extends past 1M tokens.

Nemotron 3.5 Lightning Puts 1M-Token Agents on RTX PCs
NVIDIA's Nemotron 3.5 Lightning is a 30B open model with 3B active parameters, a 1M-token context, and 4x faster output for local AI agents.

K-EXAONE 2.0 Ships 750B Open Weights Under Apache 2.0
LG AI Research released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts model with 37B active parameters, under a permissive Apache 2.0 license.

Qwen3.8-Max Packs 2.4T Parameters Into a 1M Context
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with a 1M-token window, priced at $2 per million input tokens and open weights next week.

Sparse Mixture of Experts Explained for 2026 Models
Why a 2.4T-parameter model can be cheaper than a 70B one: what active parameters mean, how routing works, and what MoE really costs to self-host in 2026.

DeepSeek V4-Flash 0731 Tops V4-Pro at a Third the Price
DeepSeek's retrained V4-Flash 0731 beats its own V4-Pro preview on every published agentic benchmark at $0.28 per million output tokens, MIT licensed.

Kimi K3 Open Weights Ship 2.8T Parameters and 1M Context
Moonshot AI released Kimi K3 open weights on July 26: 2.8 trillion parameters, 104B active per token, 1M context, and a modified MIT license.

Laguna S 2.1 Open-Weight Coding Model Fits One Desktop
Poolside's Laguna S 2.1 packs 118B parameters, activates just 8B per token, scores 70.2% on Terminal-Bench 2.1, and runs on a single desktop.

Inkling Is a 975B Open-Weights Model Under Apache 2.0
Thinking Machines released Inkling, a 975B-parameter Apache 2.0 model with 41B active, a 1M-token context window, and native four-modality reasoning.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

Kimi K3 Becomes the Largest Open-Weight AI Model Yet
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks 3rd on GDPval-AA v2, with full weights arriving July 27.

Thinking Machines Inkling: A 975B Open-Weights Model
Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model with 41B active per token and a 1M-token context window.

Kimi K2.7 Code Brings Open Weights to GitHub Copilot
GitHub Copilot's first open-weight model, Kimi K2.7 Code, reached Business and Enterprise plans on July 7 — a 1T-parameter MoE with 32B active.

LongCat-2.0: A 1.6-Trillion-Parameter Open Coding Model Hits Frontier Scores
Meituan's LongCat-2.0 is a 1.6T-parameter open-source agentic coding model matching GPT-5.5 on SWE-bench Pro, with a 1M-token context and MIT license.

Moonshot's Kimi K2.7 Code Arrives as an Efficient Open-Weight Coding Model
Moonshot AI's open-weight Kimi K2.7 Code launched June 12 with a 1T-parameter MoE design, a 256K context window, and roughly 30% lower reasoning-token use.

NVIDIA Nemotron 3 Ultra: A 550B Open Model Built for Long-Running AI Agents
NVIDIA released Nemotron 3 Ultra on June 4, 2026 — a fully open 550B-parameter reasoning model topping US open-model benchmarks and tuned for long-running AI agents.

NVIDIA Nemotron 3 Ultra: Its Most Capable Open-Weight LLM Lands at Computex 2026
NVIDIA's Nemotron 3 Ultra debuts at Computex 2026: a 550B sparse Mixture-of-Experts open-weight LLM topping the US Intelligence Index at 48, with open datasets.

Microsoft's MAI-Thinking-1: Its First In-House Reasoning Model
Microsoft's first in-house reasoning model, MAI-Thinking-1, debuts at Build 2026 with a trillion-parameter Mixture-of-Experts design and standout AIME scores.

DeepSeek V4-Pro and V4-Flash Are Here: Open-Source AI With a 1M-Token Context Window
DeepSeek drops V4-Pro (1.6T params) and V4-Flash today with 1M-token context, hybrid attention, and pricing that challenges every closed-source frontier model.

GLM-5.1 Goes Open-Source and Hits #1 on SWE-Bench Pro — Beating Every Closed AI Model
Z.ai's GLM-5.1 is a 754B open-weight MoE model under the MIT license — and it just took #1 on SWE-Bench Pro, outscoring every major closed model.

Meta Launches Llama 4 Scout and Maverick: Multimodal MoE AI Goes Open-Weight
Meta's Llama 4 Scout and Maverick bring multimodal mixture-of-experts AI to the open-source community, with an unprecedented 10 million token context window.

NVIDIA Launches Nemotron 3 Open Models at GDC — 120B Parameters With 5x the Throughput
NVIDIA's Nemotron 3 family ships in Nano, Super, and Ultra sizes with up to 1M-token context, already adopted by CrowdStrike, Cursor, Perplexity, and Zoom.

NVIDIA Debuts Nemotron 3 — Open Models With 4x Throughput and a Million-Token Context Window
NVIDIA's Nemotron 3 family ships three tiers of open models optimized for agentic AI, plus 3 trillion tokens of training data for the community.
