Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Forlinx RK1828 M.2 Card: 20 TOPS for Local LLMs on Rockchip

Forlinx RK1828 M.2 Card: 20 TOPS for Local LLMs on Rockchip

The Forlinx RK1820/RK1828 M.2 card adds 20 TOPS and up to 5GB of on-package memory to Rockchip boards, with 70 tokens/s on a 7B model per Forlinx.

Alex Circuit
Alex Circuit★Sep 29, 2026★3 min read

The Forlinx RK1828 M.2 card is one of the more practical local LLM upgrades to reach Rockchip single board computers. LinuxGizmos reported the new listing on September 28, 2026: an M.2 2280 accelerator built on Rockchip's RK1820 or RK1828 AI coprocessor, rated at 20 TOPS and designed to run 3B to 7B parameter models on one card, or much larger models across several.

  • The card uses Rockchip's RK1820 (2.5GB) or RK1828 (5GB) with DRAM stacked in the same package.
  • It is rated at 20 TOPS INT8 and supports INT4, INT8, INT16, FP8, FP16 and BF16.
  • Forlinx reports about 70 tokens per second on Qwen2.5-7B with a single RK1828 card.
  • Four cascaded cards run a 27B model at about 13 tokens per second, drawing roughly 40W in total, per Forlinx.

What is the Forlinx RK1820/RK1828 M.2 card?

It is an M.2 2280 M-key module that connects over a single PCIe 2.1 lane and pairs with Rockchip RK3568, RK3572, RK3576 and RK3588 host boards running Linux or Android. The RK182x chips are dedicated AI coprocessors; the host SoC handles the operating system and applications while the card handles inference.

The clever part is memory. Rather than streaming model weights across the slow PCIe link, the RK1828 keeps 5GB of DRAM stacked in the same package as the NPU, so a quantized 7B model can live right next to the compute. That design choice is why a single-lane link can still deliver usable token rates. Forlinx has been busy on the Rockchip side all year, including the RK3572 system-on-module it launched in June.

How fast is the RK1828 for local LLMs?

Forlinx's own figures, measured with an RK1828 on its OK3588-C development board, are:

  • Qwen2.5-3B: about 102 tokens per second
  • Qwen2.5-7B: about 70 tokens per second

Both numbers are comfortably faster than reading speed, which is what matters for a chat assistant or a voice interface on an edge device. These are vendor benchmarks, so treat them as a guide until independent testing arrives.

Can you chain multiple cards for bigger models?

Yes, and this is the headline feature. Forlinx says the cards support PCIe cascading, splitting a model's layers across multiple accelerators. With four cards, it reports running a 27B parameter model at about 13 tokens per second with 819ms to the first token, and roughly 40W of total power. It cites support for models in the 27B to 31B range in multi-card setups.

For a fanless edge AI box that needs to keep data on site, a 27B model at 40W is an attractive combination compared with adding a discrete GPU.

Software support and availability

The card uses Rockchip's RKNN3 SDK, which accepts models from TensorFlow, PyTorch and ONNX, and Forlinx lists the Qwen and Llama families among supported models. Pricing has not been announced; buyers are directed to Forlinx sales.

Why this matters for Rockchip boards

RK3588 boards already have a capable built-in NPU, as tools like Seeed's one-click AI Lab deployment have shown, but on-board NPUs top out well short of 7B models at interactive speeds. A slot-in accelerator with its own memory gives existing Rockchip hardware a clear upgrade path for local LLM work. Keep up with edge AI hardware in our mini computers section.

Sources: LinuxGizmos: Forlinx 20 TOPS M.2 AI accelerator supports PCIe cascading for local LLM inference — September 28, 2026; Forlinx: RK1820/RK1828 AI accelerator product page — September 2026.

More Mini Computers Stories