Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Needle 2 Brings 14MB Function-Calling AI to Raspberry Pi 5

Needle 2 Brings 14MB Function-Calling AI to Raspberry Pi 5

Needle 2 is a 14MB function-calling model that turns plain English into Raspberry Pi 5 actions in about 80ms on the CPU alone, with no AI HAT needed.

Alex Circuit
Alex Circuit★Sep 22, 2026★4 min read

A Tiny Model With One Job

Most talk about local AI on a Raspberry Pi focuses on chatbots and accelerator HATs. Needle 2 takes a different path. It is a 14MB function-calling model from Cactus Compute that does exactly one thing: turn a plain-English request into a structured call to a Python function. On September 22, the official Raspberry Pi blog showed it running on a Raspberry Pi 5 using only the CPU, answering in well under a tenth of a second.

  • Size: a 14MB model, with the native model session using about 28MB of RAM and the whole Python process peaking at 43 to 46.4MB
  • Speed on a Pi 5 8GB: 76 to 149ms per request, with prefill around 461 to 488 tokens per second and decode around 248 to 314 tokens per second
  • Install: one pip package, cactus-needle (version 2.0.7 in the test), Apache 2.0 licensed
  • Hardware: no AI HAT or accelerator; the demo ran on the Pi's Arm CPU under Raspberry Pi OS

What Is a Function-Calling LLM?

A function-calling model reads what you ask for and picks the right tool, filling in its arguments. It does not write essays or answer trivia. That narrow focus is what lets Needle stay so small. In the Raspberry Pi demo, the developer registered five ordinary Python functions as tools: saving a note to a file, reading the CPU temperature through vcgencmd, switching an LED, blinking it a set number of times, and taking a photo with rpicam-still.

Then the prompts rolled in. Turn the LED on took 78ms. Blink the LED 2 times took 83ms and correctly passed the count. Take a photo took 76ms. How hot is this Raspberry Pi mapped to the temperature tool in 149ms.

What Happens When You Ask It Something Off-Topic?

This is the clever part. When asked for the capital of France, Needle returned an empty list of function calls in 92ms. For a device that controls real hardware, doing nothing when a request does not match any tool is exactly the right behavior. It is a small safety property, but it is the difference between a gadget you trust and one you unplug.

Why This Matters for Raspberry Pi Projects

Voice assistants, home automation panels, robot controllers and kiosk interfaces all share the same need: map loose human language onto a fixed menu of actions. Until now, that usually meant a cloud API or a much larger local model. At under 50MB of total memory, Needle leaves nearly all of a Pi 5's RAM free for your actual application, and it keeps working offline once the weights are downloaded.

Tools are ordinary Python functions marked with a decorator, and the Raspberry Pi post notes that Needle can be fine-tuned on a laptop for your own tool set. Raspberry Pi CEO Eben Upton summed it up in the post: Needle 2 is rather excellent.

It also pairs naturally with the other edge-AI options we have covered for the board, from the BrainChip AKD1500 neuromorphic M.2 card to LLMs running on GPU-less hardware. Use a tiny tool-caller for control, and save the heavier models for when you actually need prose.

What About Needle 3?

Cactus Compute's GitHub page lists Needle 3 as the current generation, with a new attention architecture and deployable binaries from 8 to 29MB across targets including Linux on ARM64, macOS, the browser and WASI. Needle 2 remains available through a generation setting for existing projects. The Raspberry Pi walkthrough uses Needle 2, so that is the version with published Pi 5 numbers today. Browse more single-board projects in our mini computer coverage.

Sources: Raspberry Pi — September 22, 2026; Cactus Compute on GitHub — accessed September 22, 2026.

More Mini Computers Stories