Articles Tagged “Open Weight Models”
61 articles found

Liquid AI d1 Models: Open Decision AI That Answers in 8 ms
Liquid AI's open d1-3B and d1-omni-600M decision models skip token generation and answer in one pass, as fast as 8 ms on an RTX 4090 and 50 ms on Jetson.

NVIDIA Nemotron Olympiad Recipe: What the Open Release Means
NVIDIA open-sourced the Nemotron 3 recipe behind a 535.4/600 IOI 2026 run and a gold-level 30/42 IMO score, with checkpoints, data and a new benchmark.

Mistral Large 4: What a 1T Open-Weight Model Means for You
Mistral Large 4 packs 1 trillion parameters with 49B active, costs $1.36 per million input tokens, and its open weights are due by the end of October.

EmbeddingGemma 2: On-Device Search for Text, Images, Audio
EmbeddingGemma 2 is a 740M open embedding model that searches text, images, audio and video on a phone using as little as 191MB of RAM. See how it works.

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.

Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops
Ternary Bonsai 2 27B squeezes Qwen3.8-27B into 5.9GB with 1.76-bit weights and keeps 98.2% of its benchmark score. Here's how to run it locally.

Xiaomi MiMo-V2.6 Pro: What the Top Open Model Offers
Xiaomi MiMo-V2.6 Pro scores 46 on the Artificial Analysis index, the best open-weight result yet, with MIT-licensed weights and $0.87 output pricing.

Run Local LLMs on a Mac: MLX and oMLX Step-by-Step Guide
A step-by-step guide to running local LLMs on Apple Silicon with mlx-lm and oMLX: install, chat, 4-bit quantize and serve an API on port 8000.

StepFun Step 5 Preview: 600B MoE Agent Model at $1 Input
StepFun's Step 5 Preview pairs 600B parameters with 27B active and a 1M-token context at $1 per million input tokens, with open weights due October 15.

Firefox Smart Window Runs Mistral Small 4 Privately
Mozilla's Firefox Smart Window beta now runs on Mistral Small 4 in North America and France, under a zero data retention agreement with no saved chats.

Perplexity Portable Computer Runs Local AI on Windows
Perplexity's on-device agent now runs on Windows PCs with 24GB+ RTX GPUs, keeping the model, harness, orchestrator and scheduler off the cloud.

What Mistral's €3B Round Means for Open-Weight AI
Mistral closed a €3 billion Series D at over €21 billion, led by Samsung — Europe's largest tech equity round, aimed at frontier research and compute.

DeepSeek V4.1-Flash Ships 552B Open Weights Under MIT
DeepSeek V4.1-Flash landed September 10 with MIT-licensed weights, a 1M-token context, 552B parameters, and output at $0.60 per million tokens.

K2 Horizon Ships Six Open Models With Training Data
IFM's K2 Horizon releases six Apache 2.0 models from 0.9B to 375B parameters, publishing training data, code and logs alongside the weights.

Tencent Hy4 Ships 770B Open Weights Under Apache 2.0
Tencent open-sourced Hy4 preview on August 28 with 770B total parameters, 49B active per token, a 1M-token context window and Apache 2.0 weights.

IBM Granite 4.2 Brings Open Reasoning Models Local
IBM released Granite 4.2 under Apache 2.0 in 3B, 8B and 30B sizes, with a switchable thinking mode, a 512K context window and a 57.00 SWE-bench score.

Claude Desktop Now Runs Local Models Through Ollama
Ollama's new Claude Desktop integration is one toggle: local or cloud open models appear in Claude's picker, with the full agent toolset intact.

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

GLM-5.3-Flash 3-Bit Quant Runs on 128GB of Local RAM
Z.ai's 320B GLM-5.3-Flash now runs at 3-bit on 128GB of RAM via Unsloth GGUFs, retaining 82% of top-1 accuracy at under a fifth of its 650GB size.

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.

Ornith-1.5 Open Weights Score 86.1 on Terminal-Bench
Ornith-1.5 ships MIT-licensed weights from 9B to 397B, and the flagship posts 86.1 on Terminal-Bench 2.1 while a 35B MoE sibling runs far leaner.

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

North Micro Vision Packs Document AI Into 2.4B Params
Cohere Labs released North Micro Vision, a 2.4B Apache 2.0 vision model that reads full-resolution A4 pages and scores 0.921 on DocVQA on local hardware.

AI Agents Reproduced 2,226 ICML Papers in Just 19 Days
A Hugging Face community challenge used AI coding agents to audit 2,226 ICML 2026 papers, verifying claims across 6,816 public reproduction logbooks.

Gemma Downloads Top 900 Million Across Open Models
Google's Gemma open models have now passed 900 million downloads, with Gemma 4 alone contributing over 300 million since its April 2026 launch.

GLM-5.3 Posts an 84.5% CyberGym Cyber Defense Score
Z.ai's GLM-5.3 lifts CyberGym from 77.2% to 84.5% on post-training alone, and the team is holding weights back two weeks for safety hardening.

Qwen3.8-Max Open Weights Make a 2.4T Model Downloadable
Alibaba published Qwen3.8-Max open weights on August 12: 2.4 trillion parameters, 95B active per token, and a 262K context that extends past 1M tokens.

On-Device Vision AI Reads Screens in 3GB of Memory
Liquid AI's LFM2.5-VL-3B scores 69.4% average across vision benchmarks and decodes 228 tokens/second on an M5 Max, all inside roughly 3GB of memory.

Nemotron 3.5 Lightning Puts 1M-Token Agents on RTX PCs
NVIDIA's Nemotron 3.5 Lightning is a 30B open model with 3B active parameters, a 1M-token context, and 4x faster output for local AI agents.

Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU
Meta's Muse Glimmer is a 30B open-weight agentic model that compresses under 20GB at 4-bit, so a single 24GB consumer GPU can run it locally.

NVIDIA Magpie TTS Hits 12 Languages With Open Weights
NVIDIA's 364M-parameter Magpie TTS adds Arabic, Korean, and Brazilian Portuguese, and reaches 32ms time-to-first-audio on a B200 GPU with open weights.

LFM2.5-2.6B Runs Tool-Calling AI Agents in 2.5GB of RAM
Liquid AI's LFM2.5-2.6B runs full tool-calling AI agents on a phone or Raspberry Pi, hitting 220 tokens per second in under 2.5GB of memory.

K-EXAONE 2.0 Ships 750B Open Weights Under Apache 2.0
LG AI Research released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts model with 37B active parameters, under a permissive Apache 2.0 license.

Qwen3.8-Max Packs 2.4T Parameters Into a 1M Context
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with a 1M-token window, priced at $2 per million input tokens and open weights next week.

Sparse Mixture of Experts Explained for 2026 Models
Why a 2.4T-parameter model can be cheaper than a 70B one: what active parameters mean, how routing works, and what MoE really costs to self-host in 2026.

DeepSeek V4-Flash 0731 Tops V4-Pro at a Third the Price
DeepSeek's retrained V4-Flash 0731 beats its own V4-Pro preview on every published agentic benchmark at $0.28 per million output tokens, MIT licensed.

MiniMax H3 Makes 2K Video With Native Stereo Audio
MiniMax H3 generates 15-second 2K clips with native stereo sound and tops the video editing leaderboard at 1,130 Elo, priced at 0.8 yuan per second.

Kimi K3 Open Weights Ship 2.8T Parameters and 1M Context
Moonshot AI released Kimi K3 open weights on July 26: 2.8 trillion parameters, 104B active per token, 1M context, and a modified MIT license.

Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster
Liquid AI's new LFM2.5-Encoders run 8,192-token inputs on a plain CPU roughly 3.7x faster than ModernBERT-base, from just 230M parameters and open weights.

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Laguna S 2.1 Open-Weight Coding Model Fits One Desktop
Poolside's Laguna S 2.1 packs 118B parameters, activates just 8B per token, scores 70.2% on Terminal-Bench 2.1, and runs on a single desktop.

Inkling Is a 975B Open-Weights Model Under Apache 2.0
Thinking Machines released Inkling, a 975B-parameter Apache 2.0 model with 41B active, a 1M-token context window, and native four-modality reasoning.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

SAP Closes Prior Labs Deal on Tabular Foundation Models
SAP completed its Prior Labs acquisition at over 1 billion euros and will invest another 1 billion by 2030 in open tabular foundation models.

Kimi K3 Becomes the Largest Open-Weight AI Model Yet
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks 3rd on GDPval-AA v2, with full weights arriving July 27.

Thinking Machines Inkling: A 975B Open-Weights Model
Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model with 41B active per token and a 1M-token context window.

Ollama Raises $65M to Power Local Open-Source AI
Ollama closed a $65M Series B led by Theory Ventures on July 9, growing to 8.9M monthly developers and a presence in 85% of the Fortune 500.

Poolside's Laguna XS 2.1 Puts a Free Open-Weight Coding Model on Your Machine
Poolside released Laguna XS 2.1 on July 2, 2026 — a free, permissively licensed open-weight coding model that scores 70.9% on SWE-bench Verified and runs locally.

Xiaomi's HarnessX: Agents That Rewrite Their Own Scaffolding
Xiaomi's HarnessX lets AI agents rewrite their own scaffolding mid-task, delivering a +14.5% average gain, with smaller open models benefiting the most.

GLM-5.2 Open Weights Arrive as a Top Coding Model at a Fraction of the Cost
Z.ai released GLM-5.2 open weights under an MIT license on June 16, 2026 — an open-weight coding model that rivals the best closed systems on long-horizon benchmarks at roughly one-sixth the cost.

Qwen-Robot Suite Brings Open Embodied AI to Manipulation, World Modeling, and Navigation
Alibaba's Qwen team released three open embodied-AI models for robot manipulation, video world modeling, and vision-language navigation, with public weights and code.

GLM-5.2 Arrives With a Usable 1M-Token Context and MIT Open Weights
Z.ai's GLM-5.2 is a coding-first open-weight model with a usable 1-million-token context window and MIT-licensed weights that drop into agentic dev tools.

Holo3.1 Brings Fast, Private Computer-Use AI Agents to Your Own Machine
H Company's open-weight Holo3.1 agents automate desktop and mobile tasks locally, with sizes from 0.8B to 35B and quantized builds that run on consumer hardware.

MiniMax M3: An Open-Weight Model With Frontier Coding and a 1M-Token Context
MiniMax M3, released June 1, 2026, is an open-weight LLM pairing frontier-level coding, a 1-million-token context window, and native multimodality — and the weights are coming to Hugging Face.

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

HiDream-O1-Image Goes Open Source — An 8B Reasoning Image Model Lands on Hugging Face
HiDream-AI open-sourced HiDream-O1-Image on May 8, 2026 — an 8-billion parameter reasoning-driven image generation model with a Dev variant and prompt agent, free on Hugging Face.

Hugging Face Brings Open-Source LLMs to GitHub Copilot Chat in VS Code
Hugging Face wired its inference network directly into GitHub Copilot Chat on April 28, 2026 — letting VS Code developers swap in open-source LLMs from hundreds of providers right next to Copilot's default models, no extension switching required.

Mistral Forge Lets Enterprises Train Custom AI Models on Their Own Data
Mistral's new Forge platform gives enterprises end-to-end custom model training using open-weight models including the new 119B Mistral Small 4, backed by NVIDIA at GTC 2026.
