Articles Tagged “Mixture Of Experts”
35 articles found

NVIDIA Nemotron Olympiad Recipe: What the Open Release Means
NVIDIA open-sourced the Nemotron 3 recipe behind a 535.4/600 IOI 2026 run and a gold-level 30/42 IMO score, with checkpoints, data and a new benchmark.

Raspberry Pi 5 Cluster Runs Qwen3-30B at 15 Tokens a Second
Four 16GB Raspberry Pi 5 boards ran Qwen3-30B-A3B at 15.1 tokens per second on CPU alone, a 16% gain over the prior record with distributed-llama.

Mistral Large 4: What a 1T Open-Weight Model Means for You
Mistral Large 4 packs 1 trillion parameters with 49B active, costs $1.36 per million input tokens, and its open weights are due by the end of October.

Xiaomi MiMo-V2.6 Pro: What the Top Open Model Offers
Xiaomi MiMo-V2.6 Pro scores 46 on the Artificial Analysis index, the best open-weight result yet, with MIT-licensed weights and $0.87 output pricing.

StepFun Step 5 Preview: 600B MoE Agent Model at $1 Input
StepFun's Step 5 Preview pairs 600B parameters with 27B active and a 1M-token context at $1 per million input tokens, with open weights due October 15.

DeepSeek V4.1-Flash Ships 552B Open Weights Under MIT
DeepSeek V4.1-Flash landed September 10 with MIT-licensed weights, a 1M-token context, 552B parameters, and output at $0.60 per million tokens.

Qwen3.8-Max-0902 Takes Top Spot in Code Arena WebDev
Alibaba's Qwen3.8-Max-0902 debuts at 1,691 points on Code Arena WebDev and more than doubles its TerminalBench score, at unchanged $2/$6 pricing.

K2 Horizon Ships Six Open Models With Training Data
IFM's K2 Horizon releases six Apache 2.0 models from 0.9B to 375B parameters, publishing training data, code and logs alongside the weights.

Tencent Hy4 Ships 770B Open Weights Under Apache 2.0
Tencent open-sourced Hy4 preview on August 28 with 770B total parameters, 49B active per token, a 1M-token context window and Apache 2.0 weights.

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

Ornith-1.5 Open Weights Score 86.1 on Terminal-Bench
Ornith-1.5 ships MIT-licensed weights from 9B to 397B, and the flagship posts 86.1 on Terminal-Bench 2.1 while a 35B MoE sibling runs far leaner.

Qwen3.8-Max Open Weights Make a 2.4T Model Downloadable
Alibaba published Qwen3.8-Max open weights on August 12: 2.4 trillion parameters, 95B active per token, and a 262K context that extends past 1M tokens.

Nemotron 3.5 Lightning Puts 1M-Token Agents on RTX PCs
NVIDIA's Nemotron 3.5 Lightning is a 30B open model with 3B active parameters, a 1M-token context, and 4x faster output for local AI agents.

K-EXAONE 2.0 Ships 750B Open Weights Under Apache 2.0
LG AI Research released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts model with 37B active parameters, under a permissive Apache 2.0 license.

Qwen3.8-Max Packs 2.4T Parameters Into a 1M Context
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with a 1M-token window, priced at $2 per million input tokens and open weights next week.

Sparse Mixture of Experts Explained for 2026 Models
Why a 2.4T-parameter model can be cheaper than a 70B one: what active parameters mean, how routing works, and what MoE really costs to self-host in 2026.

DeepSeek V4-Flash 0731 Tops V4-Pro at a Third the Price
DeepSeek's retrained V4-Flash 0731 beats its own V4-Pro preview on every published agentic benchmark at $0.28 per million output tokens, MIT licensed.

Kimi K3 Open Weights Ship 2.8T Parameters and 1M Context
Moonshot AI released Kimi K3 open weights on July 26: 2.8 trillion parameters, 104B active per token, 1M context, and a modified MIT license.

Laguna S 2.1 Open-Weight Coding Model Fits One Desktop
Poolside's Laguna S 2.1 packs 118B parameters, activates just 8B per token, scores 70.2% on Terminal-Bench 2.1, and runs on a single desktop.

Inkling Is a 975B Open-Weights Model Under Apache 2.0
Thinking Machines released Inkling, a 975B-parameter Apache 2.0 model with 41B active, a 1M-token context window, and native four-modality reasoning.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

Kimi K3 Becomes the Largest Open-Weight AI Model Yet
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks 3rd on GDPval-AA v2, with full weights arriving July 27.

Thinking Machines Inkling: A 975B Open-Weights Model
Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model with 41B active per token and a 1M-token context window.

Kimi K2.7 Code Brings Open Weights to GitHub Copilot
GitHub Copilot's first open-weight model, Kimi K2.7 Code, reached Business and Enterprise plans on July 7 — a 1T-parameter MoE with 32B active.

LongCat-2.0: A 1.6-Trillion-Parameter Open Coding Model Hits Frontier Scores
Meituan's LongCat-2.0 is a 1.6T-parameter open-source agentic coding model matching GPT-5.5 on SWE-bench Pro, with a 1M-token context and MIT license.

Moonshot's Kimi K2.7 Code Arrives as an Efficient Open-Weight Coding Model
Moonshot AI's open-weight Kimi K2.7 Code launched June 12 with a 1T-parameter MoE design, a 256K context window, and roughly 30% lower reasoning-token use.

NVIDIA Nemotron 3 Ultra: A 550B Open Model Built for Long-Running AI Agents
NVIDIA released Nemotron 3 Ultra on June 4, 2026 — a fully open 550B-parameter reasoning model topping US open-model benchmarks and tuned for long-running AI agents.

NVIDIA Nemotron 3 Ultra: Its Most Capable Open-Weight LLM Lands at Computex 2026
NVIDIA's Nemotron 3 Ultra debuts at Computex 2026: a 550B sparse Mixture-of-Experts open-weight LLM topping the US Intelligence Index at 48, with open datasets.

Microsoft's MAI-Thinking-1: Its First In-House Reasoning Model
Microsoft's first in-house reasoning model, MAI-Thinking-1, debuts at Build 2026 with a trillion-parameter Mixture-of-Experts design and standout AIME scores.

DeepSeek V4-Pro and V4-Flash Are Here: Open-Source AI With a 1M-Token Context Window
DeepSeek drops V4-Pro (1.6T params) and V4-Flash today with 1M-token context, hybrid attention, and pricing that challenges every closed-source frontier model.

GLM-5.1 Goes Open-Source and Hits #1 on SWE-Bench Pro — Beating Every Closed AI Model
Z.ai's GLM-5.1 is a 754B open-weight MoE model under the MIT license — and it just took #1 on SWE-Bench Pro, outscoring every major closed model.

Meta Launches Llama 4 Scout and Maverick: Multimodal MoE AI Goes Open-Weight
Meta's Llama 4 Scout and Maverick bring multimodal mixture-of-experts AI to the open-source community, with an unprecedented 10 million token context window.

NVIDIA Launches Nemotron 3 Open Models at GDC — 120B Parameters With 5x the Throughput
NVIDIA's Nemotron 3 family ships in Nano, Super, and Ultra sizes with up to 1M-token context, already adopted by CrowdStrike, Cursor, Perplexity, and Zoom.

NVIDIA Debuts Nemotron 3 — Open Models With 4x Throughput and a Million-Token Context Window
NVIDIA's Nemotron 3 family ships three tiers of open models optimized for agentic AI, plus 3 trillion tokens of training data for the community.
