
Run Local LLMs on a Mac: MLX and oMLX Step-by-Step Guide
A step-by-step guide to running local LLMs on Apple Silicon with mlx-lm and oMLX: install, chat, 4-bit quantize and serve an API on port 8000.
Why Run LLMs Locally on a Mac?
Apple Silicon Macs are some of the best machines you can buy for running large language models at home. The reason is unified memory: the CPU and GPU share one large pool of fast RAM, so a Mac with 32GB or 64GB can hold models that would need an expensive graphics card anywhere else. MLX is Apple's open-source machine learning framework built for exactly that design, and it has become the fastest, most native way to run local LLMs on a Mac.
- What you need: an Apple Silicon Mac (M1 or newer) running macOS 15 Sequoia or later
- The two tools: mlx-lm for running and converting models, and oMLX for serving them as an OpenAI- and Anthropic-compatible API
- Memory rule of thumb: a 4-bit model needs a little over half a gigabyte of RAM per billion parameters, plus room for context
- Cost: both tools are free and open source (mlx-lm from Apple's ml-explore team, oMLX under Apache 2.0)
This guide walks through the whole path, from a first prompt with mlx-lm to a proper local API server with oMLX that any app can talk to. It is timely: the new Mac mini with M6 and M5 Pro chips reaches customers today, September 22, and oMLX creator Jun Kim has just joined Hugging Face to work on the project full time.
Quick Picks: Which Tool Should You Use?
- Just want to chat with a model tonight: install mlx-lm and run its built-in chat command. Five minutes, no configuration
- Want to script or build with local AI in Python: use the mlx-lm Python API with load and generate
- Want a local server for coding agents, chat apps or several models at once: install oMLX, point it at a models folder and use the endpoint on port 8000
- Want to shrink a model yourself: use mlx-lm's convert command with the quantize flag
How Much RAM Do You Need for Local LLMs on a Mac?
Memory decides what you can run, so start here. These are practical starting points for 4-bit quantized models, the sweet spot between quality and size:
- 16GB Mac: 3B to 8B parameter models run comfortably, leaving room for your other apps. Great for drafting, summarizing and quick coding help
- 24GB to 32GB Mac: 12B to 14B models run well, and 20B-class models fit with shorter context windows
- 48GB to 64GB Mac: 30B to 32B dense models, or mixture-of-experts models whose total weights fit in memory, become comfortable daily drivers
- 96GB and up (Mac Studio territory): 70B-class dense models and large MoE models with long context
Remember that context uses memory too. A long document or a big codebase in the prompt adds to the weight footprint, so leave headroom. When we tested the Mac mini M6 for local LLM prompts last month, memory bandwidth was the spec that separated the configurations most clearly.
Step 1: Install mlx-lm
mlx-lm is the official MLX package for text generation. Install it with pip, ideally inside a fresh virtual environment so it does not collide with other Python projects:
- With pip:
pip install mlx-lm - With conda:
conda install -c conda-forge mlx-lm
That is the whole install. There is no separate GPU driver to configure, because MLX talks to the Apple GPU directly.
Step 2: Generate Your First Response
Run a single prompt from the terminal:
mlx_lm.generate --prompt "How tall is Mt Everest?"
With no model specified, mlx-lm downloads its default model, a 4-bit build of Llama 3.2 3B Instruct from the mlx-community organization on Hugging Face, and answers your question. To pick a different model, add the model flag with any MLX-format repository name, such as one of the thousands of pre-converted models under mlx-community.
For a back-and-forth conversation, use the interactive chat mode instead:
mlx_lm.chat
Chat keeps the conversation history for you, which makes it the easiest way to get a feel for how a model behaves before you build anything on top of it.
Step 3: Use mlx-lm From Python
When you are ready to build something, the Python API takes a few lines. You load a model and tokenizer, wrap your message in the model's chat template, and call generate:
from mlx_lm import load, generatemodel, tokenizer = load("mlx-community/Mistral-7B-Instruct-v0.3-4bit")- Build a messages list, apply the tokenizer's chat template with the generation prompt enabled, then call
generate(model, tokenizer, prompt=prompt, verbose=True)
For chat-style apps where text should appear as it is produced, swap in stream_generate and print each chunk as it arrives. mlx-lm also supports prompt caching, which saves time when many requests share a long common prefix like a system prompt or a reference document.
Step 4: Quantize a Model Yourself
Most popular models already have MLX conversions, but converting your own is straightforward and useful for new releases. The convert command downloads the original Hugging Face weights and writes a quantized MLX version:
mlx_lm.convert --model mistralai/Mistral-7B-Instruct-v0.3 -q
The quantize flag produces a 4-bit model by default, cutting memory use to roughly a quarter of the 16-bit original with modest quality loss. Add the upload-repo flag to publish your conversion to your own Hugging Face account. If you want the theory behind the trade-offs, our LLM quantization guide comparing GGUF, AWQ and MLX goes deeper.
Step 5: Serve Models With oMLX
A terminal chat is fine for experiments. For real work, you want a local server that coding agents, chat front-ends and your own scripts can all share. That is what oMLX does. It builds on mlx-lm and mlx-vlm and adds the pieces a daily-driver server needs.
There are three ways to install it:
- Mac app: download the DMG from the oMLX GitHub releases page and drag it to Applications. It auto-updates and installs a command-line shim
- Homebrew:
brew tap jundot/omlx https://github.com/jundot/omlxand thenbrew install jundot/omlx/omlx - From source: clone the repository and run
pip install -e .with Python 3.11 to 3.13
Then start the server and point it at a folder of MLX models:
omlx serve --model-dir ~/models
By default oMLX listens on port 8000, with an admin dashboard at localhost:8000/admin and a built-in chat page at localhost:8000/admin/chat. The dashboard handles model management, monitoring and quick benchmarks.
What Makes oMLX Different From a Basic Server?
Several features are aimed squarely at people who use local AI all day:
- Tiered KV cache: recent context stays in RAM while colder context spills to your SSD, with prefix sharing so repeated prompts are not recomputed. Set the SSD location with the paged-ssd-cache-dir flag
- Continuous batching: several requests run at once, which matters when an agent fires parallel calls. Tune it with the max-concurrent-requests flag
- Multi-model serving: models load and unload automatically with least-recently-used eviction, and you can pin favorites or set per-model timeouts
- Beyond chat: vision-language models with multi-image input, plus embeddings and rerankers such as BGE-M3 and ModernBERT for local search and RAG
- Tool calling and MCP: JSON-schema-validated tool calls, with MCP server support through a config file
- Safety rails: a memory-guard option to avoid overcommitting RAM, and an API-key flag if you expose the server beyond your own machine
How Do I Connect Apps to My Local MLX Server?
oMLX exposes both OpenAI-style and Anthropic-style endpoints: chat completions and completions at /v1, a messages endpoint for Anthropic-format clients, plus embeddings, rerank and a models list. In practice that means almost any AI app or SDK that lets you change the base URL will work. Point it at localhost:8000/v1, choose a model name from the models list, and your requests stay on your own Mac.
A good first project is to connect a coding assistant or a note-taking app you already use, then compare a local model against your usual cloud model on everyday tasks. Many people find a strong 8B to 14B local model handles the majority of routine work, with the cloud reserved for the hardest problems.
Troubleshooting Common MLX Problems
- Warning that the model is larger than available memory: macOS limits how much RAM the GPU can wire by default. The mlx-lm documentation describes raising it with
sudo sysctl iogpu.wired_limit_mb=N, choosing N larger than the model but below your total RAM. Pick a smaller model if you are close to the limit - Slow first response: the first run downloads the weights, which can be several gigabytes. Later runs load from the local cache
- Model needs custom code: some architectures require the trust-remote-code option. Only enable it for repositories you trust
- Long prompts feel slow: prompt processing, not generation, is usually the bottleneck on long inputs. Prompt caching and oMLX's SSD-backed KV cache both help when you reuse context
What Is Next for MLX?
The ecosystem is moving quickly. Hugging Face announced on September 22 that Jun Kim is joining the company to maintain oMLX as a fully funded project, still under Apache 2.0, with a goal of making it faster to turn new model architectures into reference MLX implementations. For Mac owners, that should mean new open-weight models become runnable sooner after release.
Local AI on a Mac has gone from a hobbyist curiosity to a genuinely practical setup in about two years. With mlx-lm for experiments and oMLX for everyday serving, you can have a private, capable assistant running on hardware you already own this afternoon. Find more guides in our AI coverage.
Sources: mlx-lm on GitHub — accessed September 22, 2026; oMLX on GitHub — accessed September 22, 2026; Hugging Face — September 22, 2026.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
