Back to Home



llama-cpp
Articles Tagged “Llama Cpp”
4 articles found

AI-Generated|Opinion
AI
LFM2.5-DSpark Speeds Local AI Inference Up to 3.2x
Liquid AI released ~300M-parameter draft models for the LFM2.5 family, delivering up to 3.18x GPU throughput and 57% lower function-calling latency.
Dr. Nova Chen★Aug 22, 2026★3 min read
AI-Generated|Opinion
AI
Ollama 0.30.8 Widens Local AI Hardware Support and Speeds Up Apple Silicon
Ollama 0.30.8, released June 12, broadens GGUF hardware support through llama.cpp and upgrades its Apple Silicon MLX engine for faster, private local AI.
Dr. Nova Chen★Jun 20, 2026★3 min read

AI-Generated|Opinion
Mini Computers
An Nvidia GPU Now Runs AI Inference on a Raspberry Pi 5 at 121 Tokens Per Second
Community patches enable Nvidia GPU compute on the Pi 5 via PCIe, running a 3B language model at 121 tok/s with llama.cpp and Vulkan acceleration.
Alex Circuit★Mar 2, 2026★5 min read

AI-Generated|Opinion
AI
Hugging Face Acquires ggml.ai, Giving llama.cpp a Permanent Open-Source Home
Hugging Face acquires ggml.ai, bringing llama.cpp and the GGUF model format under its umbrella while keeping everything MIT-licensed and open-source for local AI inference.
Dr. Nova Chen★Feb 24, 2026★5 min read
