Articles Tagged “Qwen”
15 articles found

Raspberry Pi 5 Cluster Runs Qwen3-30B at 15 Tokens a Second
Four 16GB Raspberry Pi 5 boards ran Qwen3-30B-A3B at 15.1 tokens per second on CPU alone, a 16% gain over the prior record with distributed-llama.

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.

Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops
Ternary Bonsai 2 27B squeezes Qwen3.8-27B into 5.9GB with 1.76-bit weights and keeps 98.2% of its benchmark score. Here's how to run it locally.

Qwen3.8-Max-0902 Takes Top Spot in Code Arena WebDev
Alibaba's Qwen3.8-Max-0902 debuts at 1,691 points on Code Arena WebDev and more than doubles its TerminalBench score, at unchanged $2/$6 pricing.

Qwen3.8-Flash-Next Fits in 75GB With No GPU Needed
Qwen's 125B Flash-Next MoE runs locally in 75GB of RAM with no GPU VRAM required, scoring 62.5 on SWE-bench Pro with just 6B active parameters.

Qwen3.8-27B Runs a 262K-Context Vision Model Locally
Alibaba's Qwen3.8-27B lands under Apache 2.0 with vision, a 262K context, and a 17GB quantization that runs at 15-30 tokens per second on a laptop.

Qwen3.8-Max Open Weights Make a 2.4T Model Downloadable
Alibaba published Qwen3.8-Max open weights on August 12: 2.4 trillion parameters, 95B active per token, and a 262K context that extends past 1M tokens.

Qwen3.8-Max Packs 2.4T Parameters Into a 1M Context
Alibaba's Qwen3.8-Max is a 2.4-trillion-parameter sparse MoE model with a 1M-token window, priced at $2 per million input tokens and open weights next week.

Qwen3.7 Flash Brings 1M-Token Vision at $0.03 per Million
Alibaba's Qwen3.7 Flash is a native vision-language model with a 1M-token context window, priced at $0.03 per million input tokens with tool calling.

Qwen3.8-Max Benchmarks: What to Watch in the Preview
Alibaba's Qwen3.8-Max preview brings 2.4 trillion parameters and multimodal input, but no benchmark table yet. Here's the baseline to measure it against.

Xiaomi's HarnessX: Agents That Rewrite Their Own Scaffolding
Xiaomi's HarnessX lets AI agents rewrite their own scaffolding mid-task, delivering a +14.5% average gain, with smaller open models benefiting the most.

Qwen-AgentWorld Is an Open Model That Simulates Worlds for AI Agents
Alibaba's Qwen team open-sourced AgentWorld on June 24, 2026 — a language world model that simulates digital environments so AI agents can practice and improve.

Alibaba's Qwen3.7-Plus Pairs Vision With Autonomous Agent Skills at $0.40 per Million Tokens
Alibaba's Qwen team launched Qwen3.7-Plus on June 2, 2026 — a multimodal model combining vision and video understanding with deep reasoning, tool use, and autonomous iteration.

Alibaba's Qwen3.6-Plus Delivers 1M-Token Context and Repository-Level Agentic Coding
Qwen3.6-Plus arrives with a default 1 million-token context window and breakthrough agentic coding performance, enabling AI that can navigate and rewrite entire software repositories autonomously.
