
Raspberry Pi 5 Cluster Runs Qwen3-30B at 15 Tokens a Second
Four 16GB Raspberry Pi 5 boards ran Qwen3-30B-A3B at 15.1 tokens per second on CPU alone, a 16% gain over the prior record with distributed-llama.
A 30-billion-parameter language model running at a usable speed on hobbyist hardware sounds like a stretch, but a new benchmark shows four Raspberry Pi 5 boards doing exactly that. As reported by LinuxGizmos on October 6, 2026, Daniel Correa Villa of Hellomatik ran Qwen3-30B-A3B across a four-node Raspberry Pi 5 cluster at just over 15 tokens per second, entirely on the CPUs. It is one of the most impressive local LLM results we have seen on single board computers.
- Hardware: four Raspberry Pi 5 boards with 16GB of LPDDR4X each, NVMe SSDs, PoE+ HATs and Gigabit Ethernet.
- Model: Qwen3-30B-A3B, a mixture-of-experts model with 30B total and 3B active parameters, at Q40 quantization.
- Speed: 15.143 tokens per second for decode, and 14.449 tokens per second end to end.
- Gain: 16.1% faster than the previous public result of 13.04 tokens per second.
How Does a Raspberry Pi 5 Cluster Run Qwen3-30B?
The trick is a combination of the right model and the right software. Qwen3-30B-A3B is a mixture-of-experts design, so only about 3 billion of its parameters do work on each token. That keeps the compute per token small enough for four Cortex-A76 quad-core chips at 2.4GHz to share.
The software is a modified version of the open-source distributed-llama project, which splits the model across nodes with tensor parallelism. Each board holds part of the model in its 16GB of RAM and the boards exchange results over Gigabit Ethernet. The cluster runs Debian 13 on the 6.12 kernel.
What Changes Made It 16% Faster?
The researcher made twelve source-level changes. LinuxGizmos lists memory alignment fixes, kernel fusion, larger TCP buffers, and compiling specifically for the Cortex-A76 so the code can use the chip's NEON dot-product instructions. Together they lifted decode speed by 15.2% over unmodified distributed-llama, which managed 13.15 tokens per second on the same hardware. The modified code and benchmark scripts are on GitHub, and the full measurements are in a paper on Zenodo, first posted in May.
What Are the Limits of a Pi LLM Cluster?
The honest caveat is prompt processing. LinuxGizmos notes that a 20,000-token prompt could take more than 20 minutes before the first answer appears, so this setup suits short chats rather than long documents. The cluster is also bound by memory bandwidth, topping out at about 11.4GB/s of sustained DRAM traffic. The results have not yet been independently reproduced.
For a single-box alternative, our guide to how much memory local AI needs compares mini PC options. But as a low-power, hackable homelab project, four Pi 5 boards running a 30B model is a fantastic demonstration of how far efficient open models have come. More projects like this live in our mini computers section.
Sources: LinuxGizmos — October 6, 2026; Zenodo research paper — May 23, 2026; GitHub: hellomatik-org/distributed-llama.
More Mini Computers Stories

DSPi Firmware Turns a Raspberry Pi Pico Into a USB DSP Card
DSPi firmware turns a Raspberry Pi Pico or Pico 2 into a USB sound card with up to 110 EQ bands, active crossovers and room correction. Open source.

Raspberry Pi Desktop for PC and Mac: Revive an Old Laptop
Raspberry Pi Desktop returns for PCs and Intel Macs on 64-bit Debian Trixie, runs live from USB, and installs free on machines up to about 15 years old.

ESP32 SDR Firmware: Turn a Dev Board Into a Radio Scanner
A hidden ESP32 mode lets free ESP-SDR firmware capture raw radio from 2.2 to 2.7 GHz, and 4.8 to 6 GHz on the ESP32-C5, at up to 80 MSa/s. Here is how.
