
Local LLMs on a GPU-Less 14-Year-Old Server
A $600 secondhand Dell PowerEdge R720 with 348 GB of DDR3 and no GPU at all runs flagship open models at roughly four tokens per second.
Four Tokens a Second, No Graphics Card
The prevailing wisdom on running large models locally is that you need VRAM, and lots of it. A project written up by Hackaday this weekend takes the opposite route: a Dell PowerEdge R720 — a machine that turns fourteen this year — stuffed with 348 GB of DDR3 and not a single GPU, running GLM 5.3 Flash, Qwen 3.8 Flash and Qwen 3.8 27B entirely on its twin Xeons.
- Hardware: Dell PowerEdge R720, dual Xeon CPUs, 348 GB DDR3, no discrete graphics
- Models: GLM 5.3 Flash, Qwen 3.8 Flash and Qwen 3.8 27B, all running to completion
- Throughput: roughly four tokens per second at the top end
- Cost: about $600 on the secondhand market, per the project's own accounting
Why Does This Work at All?
Because the binding constraint for a large mixture-of-experts model is capacity before it is bandwidth. A model that will not fit simply will not run, at any speed. A retired dual-socket server is one of the cheapest ways on earth to assemble several hundred gigabytes of addressable memory, because the enterprise market dumped DDR3 registered DIMMs years ago and nobody wanted them.
What you give up is throughput. DDR3 on a Sandy Bridge-era platform delivers a small fraction of the bandwidth of modern GPU memory, and four tokens per second is roughly a slow typist. For an interactive chat session that is genuinely unpleasant. For a queue of jobs you submit and collect later, it is fine — and that framing, batch rather than interactive, is the honest case for this build.
What Is This Actually Good For?
The realistic workloads are the ones where latency does not matter and volume does. Overnight document summarisation. Bulk classification or tagging of an archive. Generating structured extractions from a few thousand files. Running an evaluation suite against a model you are considering. In each case the machine works while you are not watching, and four tokens per second across eight unattended hours is a great deal of text.
There is a second argument that has nothing to do with speed. A 2012 server that would otherwise be scrap metal becomes a functioning inference box for the price of a mid-range graphics card, and every byte of it stays on your own network. That is the same calculation that has driven the self-hosted movement all along, and it is why we keep returning to what actually fits on what — most directly in our look at how much RAM a local LLM really needs.
What Should You Check Before Buying One?
Three things, in order. First, power: a dual-socket R720 under sustained load draws real wattage continuously, and at some electricity prices the running cost overtakes the purchase price inside a year. Second, noise — 1U and 2U enterprise chassis are engineered for a machine room, not a spare bedroom. Third, memory configuration: 348 GB only helps if the DIMMs are populated across all channels on both sockets, and a cheap listing is often cheap because they are not.
If those check out, the software side is the easy part. The same runtimes that drive modern local setups work here without modification, which we walked through in our comparison of Ollama, vLLM and llama.cpp as local model servers. And if this build sounds indulgent, it is worth remembering the other end of the spectrum, where people have coaxed a local LLM onto an ESP32-S3 at nine tokens per second. More homelab and self-hosted AI hardware in our mini computer coverage.
The broader point is a cheerful one. The hardware floor for running capable open models keeps dropping, and it is dropping in two directions at once — newer silicon getting more efficient, and older silicon getting cheap enough that inefficiency stops mattering.
Sources: Hackaday — September 20, 2026; MattMo's project video — September 2026.
More Mini Computers Stories

MaTouch ESP32-S3 E Ink Board: A $60 Voice-Ready Dashboard
Makerfabs' MaTouch ESP32-S3 pairs a 7.5-inch four-color E Ink screen with a four-mic array for $59.80. Here are the specs and what you can build.

ASRock NUC 300 Mini PCs: Wildcat Lake With Dual 2.5GbE
ASRock Industrial's NUC 300 series brings Intel Core 5 320 and Core 3 304 to mini PCs and boards with up to 64GB of DDR5-6400 memory and dual 2.5GbE.

DGX Spark 64GB: Who NVIDIA's $4,999 Local AI Box Is For
NVIDIA's DGX Spark 64GB starts at $4,999 on Oct 23 and runs models up to 100B parameters. Here is who the local AI box suits and how clustering works.
