Articles Tagged “Llama Cpp”
5 articles found

Transformers Now Runs GGUF Quants on Mac at llama.cpp Speed
Hugging Face Transformers can now run GGUF quants packed on Apple Silicon. A Qwen3.5-4B Q4_K_M model shrinks from 8.42GB to 2.74GB. Here's how it works.

LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.
Ollama 0.30.8 Widens Local AI Hardware Support and Speeds Up Apple Silicon
Ollama 0.30.8, released June 12, broadens GGUF hardware support through llama.cpp and upgrades its Apple Silicon MLX engine for faster, private local AI.

An Nvidia GPU Now Runs AI Inference on a Raspberry Pi 5 at 121 Tokens Per Second
Community patches enable Nvidia GPU compute on the Pi 5 via PCIe, running a 3B language model at 121 tok/s with llama.cpp and Vulkan acceleration.

Hugging Face Acquires ggml.ai, Giving llama.cpp a Permanent Open-Source Home
Hugging Face acquires ggml.ai, bringing llama.cpp and the GGUF model format under its umbrella while keeping everything MIT-licensed and open-source for local AI inference.
