
Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed
Hugging Face Transformers can now run GGUF quants packed on Apple Silicon. A Qwen3.5-4B Q4_K_M model shrinks from 8.42GB to 2.74GB. Here's how it works.
Hugging Face Transformers just learned to speak GGUF natively. In a September 22, 2026 update, the Hugging Face team showed Transformers loading llama.cpp's quantized GGUF files and running them in their packed, compressed form on Apple Silicon, instead of expanding them back to full size first. For Mac owners who build with Python, that means the same small model files people already use in llama.cpp and Ollama now work inside the most popular machine learning library, at speeds Hugging Face describes as close to llama.cpp.
- What changed: Transformers runs GGUF quants packed, using ggml kernels for quantization, normalization and attention
- Size win: Unsloth's Qwen3.5-4B drops from 8.42GB in BF16 to 2.74GB at Q4_K_M
- Serving: the transformers serve command exposes an OpenAI-compatible API on localhost port 8000
- Current limits: packed inference is Apple Silicon only, and only Qwen3.5 dense and MoE models are supported so far
What Are GGUF Quants, and Why Does This Matter?
GGUF is the file format behind llama.cpp. It stores a model's weights at reduced precision, such as 4-bit or 6-bit, so a model fits in far less memory. Until now, loading a GGUF file in Transformers meant dequantizing it, which inflated it back to full size and threw away the memory savings. The new path keeps the weights packed and runs the math directly on them with ggml kernels, the same low-level code that powers llama.cpp.
Hugging Face's example uses Unsloth's Qwen3.5-4B quants. The sizes tell the story: 8.42GB in BF16, 3.53GB at Q6_K, 3.14GB at Q5_K_M and 2.74GB at Q4_K_M, which the team suggests as the starting point. If you want a refresher on how those quant levels trade size against accuracy, our LLM quantization guide walks through GGUF, AWQ and MLX side by side.
How to Run a GGUF Model in Transformers
The workflow stays familiar. You load the model with the usual from_pretrained call and pass the name of the GGUF file as an extra argument. For now you need Transformers installed from the main branch plus Hugging Face's kernels package. From there, the transformers serve command starts a local server with an OpenAI-compatible API, so existing chat apps and agent tools can point at it without code changes.
Performance on an M2 Max MacBook Pro
Hugging Face tested on an M2 Max MacBook Pro with 32GB of memory and reports that Transformers comes close to llama.cpp across dense, larger dense and mixture-of-experts checkpoints. It is candid about the gaps: padded and batched generation is slower for now, other hardware can only dequantize, and model support starts with the Qwen3.5 family.
Who Should Try This on Their Mac?
This matters most for developers who live in Python. Tools like llama.cpp and Ollama are excellent for running a model, but research code, fine-tuning scripts and evaluation harnesses are usually written against Transformers. Being able to run the same compact GGUF file in both places removes a conversion step and keeps experiments consistent.
It also continues a clear direction. Hugging Face brought ggml and llama.cpp under its roof earlier this year, as we covered in Hugging Face's ggml.ai acquisition, and this update is a concrete payoff: the two ecosystems now share one model format. For more ways to run models locally, see our AI coverage hub.
Sources: Hugging Face — September 22, 2026; Unsloth Qwen3.5-4B GGUF — model card.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
