Articles Tagged “Open Weight Models”
41 articles found

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.

Ornith-1.5 Open Weights Score 86.1 on Terminal-Bench
Ornith-1.5 ships MIT-licensed weights from 9B to 397B, and the flagship posts 86.1 on Terminal-Bench 2.1 while a 35B MoE sibling runs far leaner.

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

North Micro Vision Packs Document AI Into 2.4B Params
Cohere Labs released North Micro Vision, a 2.4B Apache 2.0 vision model that reads full-resolution A4 pages and scores 0.921 on DocVQA on local hardware.

AI Agents Reproduced 2,226 ICML Papers in Just 19 Days
A Hugging Face community challenge used AI coding agents to audit 2,226 ICML 2026 papers, verifying claims across 6,816 public reproduction logbooks.

Gemma Downloads Top 900 Million Across Open Models
Google's Gemma open models have now passed 900 million downloads, with Gemma 4 alone contributing over 300 million since its April 2026 launch.

GLM-5.3 Posts an 84.5% CyberGym Cyber Defense Score
Z.ai's GLM-5.3 lifts CyberGym from 77.2% to 84.5% on post-training alone, and the team is holding weights back two weeks for safety hardening.

Qwen3.8-Max Open Weights Make a 2.4T Model Downloadable
Alibaba published Qwen3.8-Max open weights on August 12: 2.4 trillion parameters, 95B active per token, and a 262K context that extends past 1M tokens.

On-Device Vision AI Reads Screens in 3GB of Memory
Liquid AI's LFM2.5-VL-3B scores 69.4% average across vision benchmarks and decodes 228 tokens/second on an M5 Max, all inside roughly 3GB of memory.

Nemotron 3.5 Lightning Puts 1M-Token Agents on RTX PCs
NVIDIA's Nemotron 3.5 Lightning is a 30B open model with 3B active parameters, a 1M-token context, and 4x faster output for local AI agents.

Muse Glimmer Runs a 30B Agentic Model on One Consumer GPU
Meta's Muse Glimmer is a 30B open-weight agentic model that compresses under 20GB at 4-bit, so a single 24GB consumer GPU can run it locally.

NVIDIA Magpie TTS Hits 12 Languages With Open Weights
NVIDIA's 364M-parameter Magpie TTS adds Arabic, Korean, and Brazilian Portuguese, and reaches 32ms time-to-first-audio on a B200 GPU with open weights.

LFM2.5-2.6B Runs Tool-Calling AI Agents in 2.5GB of RAM
Liquid AI's LFM2.5-2.6B runs full tool-calling AI agents on a phone or Raspberry Pi, hitting 220 tokens per second in under 2.5GB of memory.

K-EXAONE 2.0 Ships 750B Open Weights Under Apache 2.0
LG AI Research released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts model with 37B active parameters, under a permissive Apache 2.0 license.

Qwen3.8-Max Packs 2.4T Parameters Into a 1M Context
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with a 1M-token window, priced at $2 per million input tokens and open weights next week.

Sparse Mixture of Experts Explained for 2026 Models
Why a 2.4T-parameter model can be cheaper than a 70B one: what active parameters mean, how routing works, and what MoE really costs to self-host in 2026.

DeepSeek V4-Flash 0731 Tops V4-Pro at a Third the Price
DeepSeek's retrained V4-Flash 0731 beats its own V4-Pro preview on every published agentic benchmark at $0.28 per million output tokens, MIT licensed.

MiniMax H3 Makes 2K Video With Native Stereo Audio
MiniMax H3 generates 15-second 2K clips with native stereo sound and tops the video editing leaderboard at 1,130 Elo, priced at 0.8 yuan per second.

Kimi K3 Open Weights Ship 2.8T Parameters and 1M Context
Moonshot AI released Kimi K3 open weights on July 26: 2.8 trillion parameters, 104B active per token, 1M context, and a modified MIT license.

Liquid AI Encoders Hit 8K Context on CPU 3.7x Faster
Liquid AI's new LFM2.5-Encoders run 8,192-token inputs on a plain CPU roughly 3.7x faster than ModernBERT-base, from just 230M parameters and open weights.

LLM Quantization Guide: GGUF vs AWQ vs MLX in 2026
A practical guide to LLM quantization formats — GGUF, AWQ, GPTQ and MLX — with VRAM math, quality trade-offs and picks for every kind of machine.

Laguna S 2.1 Open-Weight Coding Model Fits One Desktop
Poolside's Laguna S 2.1 packs 118B parameters, activates just 8B per token, scores 70.2% on Terminal-Bench 2.1, and runs on a single desktop.

Inkling Is a 975B Open-Weights Model Under Apache 2.0
Thinking Machines released Inkling, a 975B-parameter Apache 2.0 model with 41B active, a 1M-token context window, and native four-modality reasoning.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

SAP Closes Prior Labs Deal on Tabular Foundation Models
SAP completed its Prior Labs acquisition at over 1 billion euros and will invest another 1 billion by 2030 in open tabular foundation models.

Kimi K3 Becomes the Largest Open-Weight AI Model Yet
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that ranks 3rd on GDPval-AA v2, with full weights arriving July 27.

Thinking Machines Inkling: A 975B Open-Weights Model
Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model with 41B active per token and a 1M-token context window.

Ollama Raises $65M to Power Local Open-Source AI
Ollama closed a $65M Series B led by Theory Ventures on July 9, growing to 8.9M monthly developers and a presence in 85% of the Fortune 500.

Poolside's Laguna XS 2.1 Puts a Free Open-Weight Coding Model on Your Machine
Poolside released Laguna XS 2.1 on July 2, 2026 — a free, permissively licensed open-weight coding model that scores 70.9% on SWE-bench Verified and runs locally.

Xiaomi's HarnessX: Agents That Rewrite Their Own Scaffolding
Xiaomi's HarnessX lets AI agents rewrite their own scaffolding mid-task, delivering a +14.5% average gain, with smaller open models benefiting the most.

GLM-5.2 Open Weights Arrive as a Top Coding Model at a Fraction of the Cost
Z.ai released GLM-5.2 open weights under an MIT license on June 16, 2026 — an open-weight coding model that rivals the best closed systems on long-horizon benchmarks at roughly one-sixth the cost.

Qwen-Robot Suite Brings Open Embodied AI to Manipulation, World Modeling, and Navigation
Alibaba's Qwen team released three open embodied-AI models for robot manipulation, video world modeling, and vision-language navigation, with public weights and code.

GLM-5.2 Arrives With a Usable 1M-Token Context and MIT Open Weights
Z.ai's GLM-5.2 is a coding-first open-weight model with a usable 1-million-token context window and MIT-licensed weights that drop into agentic dev tools.

Holo3.1 Brings Fast, Private Computer-Use AI Agents to Your Own Machine
H Company's open-weight Holo3.1 agents automate desktop and mobile tasks locally, with sizes from 0.8B to 35B and quantized builds that run on consumer hardware.

MiniMax M3: An Open-Weight Model With Frontier Coding and a 1M-Token Context
MiniMax M3, released June 1, 2026, is an open-weight LLM pairing frontier-level coding, a 1-million-token context window, and native multimodality — and the weights are coming to Hugging Face.

Gemma 4 12B Brings Full Multimodal AI to a 16GB Laptop — Free Under Apache 2.0
Google DeepMind released Gemma 4 12B on June 3, 2026 — an open-weight, encoder-free multimodal model with native audio that runs locally on a 16GB consumer laptop.

Ollama v0.24 Lands With Qwen 3.6 Support — Local AI Just Got a Major Upgrade for Self-Hosted LLM Builders
Ollama released v0.24.0 on May 14, 2026 with first-class support for Qwen 3.6 — bringing Alibaba's 35B-A3B mixture-of-experts model to anyone running local LLMs on their own hardware.

Google Drops Multi-Token Prediction Drafters for Gemma 4 — Up to 3x Faster Local LLM Inference With Zero Quality Loss
On May 5, 2026 Google released open Multi-Token Prediction drafters for the Gemma 4 family, delivering up to 3x faster local LLM inference without any quality loss — Apache 2.0 licensed.

HiDream-O1-Image Goes Open Source — An 8B Reasoning Image Model Lands on Hugging Face
HiDream-AI open-sourced HiDream-O1-Image on May 8, 2026 — an 8-billion parameter reasoning-driven image generation model with a Dev variant and prompt agent, free on Hugging Face.

Hugging Face Brings Open-Source LLMs to GitHub Copilot Chat in VS Code
Hugging Face wired its inference network directly into GitHub Copilot Chat on April 28, 2026 — letting VS Code developers swap in open-source LLMs from hundreds of providers right next to Copilot's default models, no extension switching required.

Mistral Forge Lets Enterprises Train Custom AI Models on Their Own Data
Mistral's new Forge platform gives enterprises end-to-end custom model training using open-weight models including the new 119B Mistral Small 4, backed by NVIDIA at GTC 2026.
