
Needle 2 Brings 14MB Function-Calling AI to Raspberry Pi 5
Needle 2 is a 14MB function-calling model that turns plain English into Raspberry Pi 5 actions in about 80ms on the CPU alone, with no AI HAT needed.
A Tiny Model With One Job
Most talk about local AI on a Raspberry Pi focuses on chatbots and accelerator HATs. Needle 2 takes a different path. It is a 14MB function-calling model from Cactus Compute that does exactly one thing: turn a plain-English request into a structured call to a Python function. On September 22, the official Raspberry Pi blog showed it running on a Raspberry Pi 5 using only the CPU, answering in well under a tenth of a second.
- Size: a 14MB model, with the native model session using about 28MB of RAM and the whole Python process peaking at 43 to 46.4MB
- Speed on a Pi 5 8GB: 76 to 149ms per request, with prefill around 461 to 488 tokens per second and decode around 248 to 314 tokens per second
- Install: one pip package, cactus-needle (version 2.0.7 in the test), Apache 2.0 licensed
- Hardware: no AI HAT or accelerator; the demo ran on the Pi's Arm CPU under Raspberry Pi OS
What Is a Function-Calling LLM?
A function-calling model reads what you ask for and picks the right tool, filling in its arguments. It does not write essays or answer trivia. That narrow focus is what lets Needle stay so small. In the Raspberry Pi demo, the developer registered five ordinary Python functions as tools: saving a note to a file, reading the CPU temperature through vcgencmd, switching an LED, blinking it a set number of times, and taking a photo with rpicam-still.
Then the prompts rolled in. Turn the LED on took 78ms. Blink the LED 2 times took 83ms and correctly passed the count. Take a photo took 76ms. How hot is this Raspberry Pi mapped to the temperature tool in 149ms.
What Happens When You Ask It Something Off-Topic?
This is the clever part. When asked for the capital of France, Needle returned an empty list of function calls in 92ms. For a device that controls real hardware, doing nothing when a request does not match any tool is exactly the right behavior. It is a small safety property, but it is the difference between a gadget you trust and one you unplug.
Why This Matters for Raspberry Pi Projects
Voice assistants, home automation panels, robot controllers and kiosk interfaces all share the same need: map loose human language onto a fixed menu of actions. Until now, that usually meant a cloud API or a much larger local model. At under 50MB of total memory, Needle leaves nearly all of a Pi 5's RAM free for your actual application, and it keeps working offline once the weights are downloaded.
Tools are ordinary Python functions marked with a decorator, and the Raspberry Pi post notes that Needle can be fine-tuned on a laptop for your own tool set. Raspberry Pi CEO Eben Upton summed it up in the post: Needle 2 is rather excellent.
It also pairs naturally with the other edge-AI options we have covered for the board, from the BrainChip AKD1500 neuromorphic M.2 card to LLMs running on GPU-less hardware. Use a tiny tool-caller for control, and save the heavier models for when you actually need prose.
What About Needle 3?
Cactus Compute's GitHub page lists Needle 3 as the current generation, with a new attention architecture and deployable binaries from 8 to 29MB across targets including Linux on ARM64, macOS, the browser and WASI. Needle 2 remains available through a generation setting for existing projects. The Raspberry Pi walkthrough uses Needle 2, so that is the version with published Pi 5 numbers today. Browse more single-board projects in our mini computer coverage.
Sources: Raspberry Pi — September 22, 2026; Cactus Compute on GitHub — accessed September 22, 2026.
More Mini Computers Stories

MaTouch ESP32-S3 E Ink Board: A $60 Voice-Ready Dashboard
Makerfabs' MaTouch ESP32-S3 pairs a 7.5-inch four-color E Ink screen with a four-mic array for $59.80. Here are the specs and what you can build.

ASRock NUC 300 Mini PCs: Wildcat Lake With Dual 2.5GbE
ASRock Industrial's NUC 300 series brings Intel Core 5 320 and Core 3 304 to mini PCs and boards with up to 64GB of DDR5-6400 memory and dual 2.5GbE.

DGX Spark 64GB: Who NVIDIA's $4,999 Local AI Box Is For
NVIDIA's DGX Spark 64GB starts at $4,999 on Oct 23 and runs models up to 100B parameters. Here is who the local AI box suits and how clustering works.
