Home›Guides›How to set up MLX on a Mac with mlx-lm, and what it sends out

Setup · Guide

How to set up MLX on a Mac with mlx-lm, and what it sends out

Install mlx-lm with pip in a native Python 3.11 or newer on an Apple silicon Mac, then run mlx_lm.chat: the first run downloads a 1.7 GiB model from Hugging Face and gives you a prompt. That is the whole setup, 2 steps. Of the 5 runtimes we compare, MLX is the only one that neither keeps your chats anywhere nor makes a single call of its own in a whole day. Its server is the other story: out of the box, any website you open can use it.

Measured and written by Local AI ScopePublished Figures measured by Local AI Scope

Install mlx-lm with pip in a native Python 3.11 or newer on an Apple silicon Mac, then run mlx_lm.chat: the first run downloads a 1.7 GiB model from Hugging Face and gives you a prompt. That is the whole setup, 2 steps. Of the 5 runtimes we compare, MLX is the only one that neither keeps your chats anywhere nor makes a single call of its own in a whole day. Its server is the other story: out of the box, any website you open can use it.

And the package has just woken up. pip install mlx-lm gives you 0.32.0, published on 1 October 2026 after more than 5 months without a release, and the mlx framework underneath is at 0.32.3, from 29 September 2026. Both versions were confirmed on PyPI on 1 October 2026, and every command and default on this page was checked against the 0.32.0 code that same day.

What MLX is, and what mlx-lm adds

MLX is Apple’s array framework for Apple silicon, the maths layer in the same family as NumPy or PyTorch, published by Apple under the MIT licence. It uses the Mac’s unified memory: CPU and GPU work on the same pool, with no copying between them. On its own it does not chat with anyone. mlx-lm is the package on top: it fetches language models from Hugging Face, generates text, serves an API and converts models. When someone says they run a model “in MLX”, the program running is mlx-lm, and it installs MLX for you.

What it is not, because each of these gets assumed:

  • Not the Neural Engine. The device types in MLX’s code are exactly 2, CPU and GPU. The Neural Engine in your chip sits idle.
  • Not a GGUF runner. It runs MLX-format and Hugging Face safetensors files. A GGUF you already have cannot be loaded, and mlx_lm.convert starts from the original Hugging Face weights, not from a GGUF.
  • Not an app. There is no window: mlx-lm installs command-line tools and a Python library. If you want MLX models behind a window, LM Studio runs them through its own MLX engine, and How to set up LM Studio and its local server, and what it sends out covers that route.

What you need: Apple silicon, macOS 14 and Python 3.11

Requirement What it means in practice
Apple silicon Intel Macs are out: every macOS build of mlx 0.32.3 on PyPI is arm64, and there is no x86_64 one
macOS 14.0 or newer macOS 15 or newer for the part of mlx-lm that speeds up large models by wiring their memory (see the memory section)
A native Python 3.11 or newer mlx-lm 0.32.0 declares Python 3.11; mlx alone would accept 3.10. On a 3.10, pip does not fail: it quietly installs the old 0.31.3, the last version that accepts it

Requirements from MLX’s installation guide at version 0.32.2, read on 27 September 2026, and from the PyPI metadata and builds of mlx 0.32.3 and mlx-lm 0.32.0, read on 1 October 2026.

A Python running under Rosetta installs nothing: pip reports that it cannot find a matching distribution. MLX’s own guide gives the check:

python -c "import platform; print(platform.processor())"

It has to print arm. If it prints i386 on an M-series Mac, that Python is not native.

MLX on Windows or Linux? The mlx framework publishes builds for both: Windows x64 and ARM64 since July 2026, and Linux with a CUDA backend (pip install "mlx[cuda12]", NVIDIA architecture SM 7.5 or newer) or CPU only (pip install "mlx[cpu]"). mlx-lm is another matter. It declares its dependency on mlx for macOS only, so a plain pip install mlx-lm on Windows or Linux installs the tools without the engine. On Linux, mlx-lm 0.32.0 offers cuda12, cuda13 and cpu extras that bring the matching mlx build. We installed the cpu one on a Linux x86_64 machine on 1 October 2026, and its server starts and answers; how well a model runs there, we have not measured. Neither package’s documentation mentions Windows.

How to install mlx-lm, and check which version you got

A virtual environment keeps it away from any other Python on the Mac:

python3 -m venv ~/mlx-env
source ~/mlx-env/bin/activate
pip install mlx-lm

The python3 in the first line has to be the native 3.11 or newer from the previous section. With conda, the line is conda install -c conda-forge mlx-lm, but on 1 October 2026 conda-forge was still serving 0.31.3.

That single install brings mlx-lm 0.32.0, mlx 0.32.3 (0.32.0 asks for mlx 0.32.2 or newer, so pip takes the latest) and transformers 5.7 or newer for the tokenizers. It also puts 18 commands on your PATH, all starting with mlx_lm: the ones this page uses are mlx_lm.chat, mlx_lm.generate, mlx_lm.server, mlx_lm.manage and mlx_lm.convert. Each one lists its options with -h.

To see what you have:

pip show mlx mlx-lm
python -c "import mlx.core as mx; print(mx.__version__)"
mlx_lm --version

The first shows both packages, the second the framework alone, the third mlx-lm alone. On 1 October 2026 the answers are 0.32.3 for mlx and 0.32.0 for mlx-lm. If mlx-lm says 0.31.3, look at your Python: on a 3.10, that is as far as pip will go.

0.32.0, the end of a long silence. mlx-lm published 19 versions between 25 August 2025 and 22 April 2026, the last of them 0.31.3, and then nothing for more than 5 months. The repository did not stop: between 22 April and 30 September 2026 it took 134 commits on 50 different days, and all that time pip kept handing out April’s code. 0.32.0, published on 1 October 2026, brings that work into the package: the folder of model definitions grows from 119 to 135 files, with 16 new ones such as mistral4, kimi_k3 and olmo_hybrid, and none removed. If you installed during the wait, one line brings you up to date: pip install -U mlx-lm, and then pip show mlx mlx-lm should say 0.32.3 for mlx and 0.32.0 for mlx-lm.

Your first model: mlx_lm.chat and mlx_lm.generate

mlx_lm.chat

Without --model, both mlx_lm.chat and mlx_lm.generate use mlx-community/Llama-3.2-3B-Instruct-4bit, whose files weigh 1.70 GiB on Hugging Face. The first run downloads it; after that it loads from disk. In the chat, q quits, r resets the conversation and h shows those commands. The conversation lives in memory while the chat is open and is written nowhere: quit and it is gone.

To pick the model, give its Hugging Face name or a local folder. The mlx-community organisation on Hugging Face holds the ready-converted ones:

mlx_lm.chat --model mlx-community/Qwen3-8B-4bit
mlx_lm.generate --model mlx-community/Qwen3-8B-4bit --prompt "How tall is Mt Everest?" --max-tokens 300

Watch --max-tokens: by default an answer stops at 100 tokens in mlx_lm.generate, 256 in mlx_lm.chat and 512 in the server. After the answer, mlx_lm.generate prints its figures: Prompt: N tokens, X tokens-per-sec, Generation: N tokens, X tokens-per-sec and Peak memory: X GB. That last line is the most honest fit test you have: what the model and the conversation actually took on your machine. Since 0.32.0, mlx_lm.chat gives you the same after every answer, in one line that ends in peak X GB.

Which model fits your Mac, and how MLX uses the memory

Unified memory is shared with macOS and everything you have open. Our hardware sheets keep 3 GiB aside for that and compare the rest against the 19 models we measure, at Q4_K_M, our quality floor, with an 8,192-token context:

Mac Left for the model Models that fit Largest that fits
Mac with M4 and 16 GB 13.00 GiB 8 of 19 gpt-oss-20b
Mac mini M6 32 GB 29.00 GiB 15 of 19 qwen3.6-35b-a3b
Mac with M4 Max and 64 GB 61.00 GiB 16 of 19 gpt-oss-120b

Figures served by each hardware sheet on 27 September 2026, from the machine data of 23 September. The model finder does the same sum for your own machine, context and language, in your browser.

Those sizes are GGUF files. An MLX 4-bit file is a different file with a different compression, so we compared the 3 that overlap with what the 16 GB Mac takes. On Hugging Face on 27 September 2026, mlx-community/Qwen3-4B-Instruct-2507-4bit weighed 2.11 GiB against 2.33 GiB at Q4_K_M, mlx-community/Qwen3-8B-4bit 4.29 GiB against 4.68, and mlx-community/gpt-oss-20b-MXFP4-Q4 10.41 GiB against 10.83. For memory, the finder’s verdict is on the safe side for all 3. For quality it says nothing: 4-bit MLX and Q4_K_M are not the same compression.

What MLX does with the memory. mlx_lm.generate and mlx_lm.server ask macOS how much memory it recommends for the GPU and wire up to that amount: mlx-lm’s way of keeping a large model fast. When a model takes more than 90% of that recommendation, mlx_lm.generate warns you: [WARNING] Generating with a model that requires … MB which is close to the maximum recommended size of … MB. This can be slow. The remedy mlx-lm documents needs macOS 15 or newer:

sudo sysctl iogpu.wired_limit_mb=N

N has to be larger than the model in megabytes and smaller than the Mac’s memory. For long conversations, --max-kv-size caps the memory the context takes with a rotating cache: 512 uses very little and answers worse, 4,096 or more uses more and answers better. And 0.32.0 brings --prefill-step-size, which the server already had, to mlx_lm.generate and mlx_lm.chat: a long prompt is read in steps of 2,048 tokens by default, and a smaller step lowers the peak while it is being read, at the cost of reading it more slowly.

How to start mlx_lm.server, the OpenAI-compatible server on port 8080

mlx_lm.server --model mlx-community/Qwen3-8B-4bit

It listens on 127.0.0.1:8080; --host and --port change that. It answers POST /v1/chat/completions and POST /v1/completions, plus GET /v1/models and GET /health. 2 checks:

curl http://127.0.0.1:8080/health
curl http://127.0.0.1:8080/v1/models

The first answers {"status": "ok"}; since 0.32.0, if the thread that generates text has died, it answers {"status": "unavailable"} with a 503 instead. The second lists every model in your Hugging Face cache that has the files mlx-lm looks for, not only the loaded one: an original you downloaded to convert shows up too. For an OpenAI client, the base URL is http://127.0.0.1:8080/v1; the server checks no key, so any string does. Unless the client sets them, answers use temperature 0.0 and 512 tokens at most.

The model field is an order. If a request names another model, the server unloads the current one and loads that one, downloading it from Hugging Face first if it is not in your cache. One model sits in memory at a time. Up to 0.31.3 there was a trap on top: if the model’s configuration named a Python file of its own, mlx-lm ran that file while loading it, without asking. 0.32.0 refuses unless the server was started with --trust-remote-code; we tried it with a test model, and 0.31.3 ran the file while 0.32.0 stopped with an error.

What the defaults leave open. There is no authentication at all, and CORS is open to any origin, any method and any header. The address keeps other machines out; the CORS setting lets in any web page you open in your browser, and that page can read the answers and, through the model field, make your Mac download a model. mlx-lm’s documentation is plain about the server: The MLX LM server is not recommended for production as it only implements basic security checks. The server prints the same warning every time it starts. It is built on Python’s standard HTTP server. What closes most of it:

  • Leave --host at 127.0.0.1.
  • Name the pages that may use it: --allowed-origins http://localhost:3000, a comma-separated list, instead of the default *.
  • If another machine needs it, reach it through an SSH tunnel (ssh -L 8080:127.0.0.1:8080 you@your-mac) rather than --host 0.0.0.0.

Where MLX keeps models, and how to delete them

What Where
Downloaded models The Hugging Face cache: ~/.cache/huggingface/hub, movable with HF_HOME or HF_HUB_CACHE
Models you convert ./mlx_model in the folder you ran the command from, unless you pass --mlx-path
Chats Nowhere. Nothing is written
Calibration text for AWQ, GPTQ or dynamic quantisation ~/.cache/mlx-lm/, created only if you use those methods

Paths from mlx-lm’s documentation and its code, and from Hugging Face’s documentation of its cache variables, read on 27 September 2026; mlx-lm’s checked again against the 0.32.0 code on 1 October 2026.

The cache is Hugging Face’s, not MLX’s, so other Hugging Face tools on the Mac share it. mlx-lm manages it with one command:

mlx_lm.manage --scan
mlx_lm.manage --delete --pattern mlx-community/Qwen3-8B-4bit

--scan only lists repositories whose name contains mlx unless you give another --pattern, so a model you loaded straight from its original repository will not appear in the plain scan.

Converting and quantising a model with mlx_lm.convert

When the model you want is not in mlx-community, convert it yourself:

mlx_lm.convert --model Qwen/Qwen3-8B -q

-q on its own means 4 bits in groups of 64, in the affine mode. --q-bits 8 changes the bits; --q-mode also takes mxfp4, nvfp4 and mxfp8; --quant-predicate takes 4 mixed-bit recipes, from mixed_2_6 to mixed_4_6. The result goes to ./mlx_model, and the command refuses to run if that folder already exists.

Count the download before you start. Conversion begins from the original weights: Qwen3-8B has 8.2 billion parameters in BF16, about 15.3 GiB to fetch for a result of about 4.3 GiB. --upload-repo would publish the converted model to your Hugging Face account; without it, nothing leaves.

What MLX sends out on its own, and what leaves when you ask

On its own, nothing. The MLX (mlx-lm): what it sends home and whether you can audit it sheet counts 0 calls in an idle day: no update check, no account, no telemetry of its own. On 25 August 2026 we swept the whole mlx_lm package, both the main branch and the published 0.31.3, for telemetry and analytics code and found none, and the same in the mlx framework. On 1 October 2026 we repeated the sweep on the published 0.32.0: still none.

The network starts with your first command. mlx-lm downloads through Hugging Face’s own library, huggingface_hub, from huggingface.co, and not only the first time: each time you load a model by its Hugging Face name, that library asks huggingface.co whether the files have a newer version, even when they are already cached. What goes with the request is a user-agent header with your Python version, the huggingface_hub version and, if you have it installed, PyTorch’s. mlx-lm does not identify itself, so the header literally starts with unknown/None. Hugging Face’s documentation says its libraries collect some usage data by default, and gives the switch.

Once your models are downloaded, 2 variables close it:

export HF_HUB_OFFLINE=1
export HF_HUB_DISABLE_TELEMETRY=1
mlx_lm.chat --model mlx-community/Qwen3-8B-4bit

With HF_HUB_OFFLINE=1 no HTTP call is made to Hugging Face: a cached model loads, and a model that is not cached fails with an error instead of downloading. DO_NOT_TRACK=1 works like the second line.

Everything else leaves only when you ask for it: --upload-repo and mlx_lm.upload publish to Hugging Face, AWQ, GPTQ and dynamic quantisation fetch a calibration text from a GitHub gist once, DWQ downloads a Hugging Face dataset, fine-tuning with a Hugging Face dataset name downloads the dataset, and the server’s model field downloads whatever a client names. With MLXLM_USE_MODELSCOPE=true, downloads go to ModelScope instead of Hugging Face.

The objection is fair. “No telemetry of its own” is a statement about mlx-lm’s code, not about your network. It was read from the published source, not from a traffic capture, and the sheet says so. The one library that talks, huggingface_hub, is not Apple’s, and its version is whatever pip resolved on your machine. Even so, it is a count read from the code, and that is already more than LM Studio allows: its app is closed source, so what it sends can only be described from what its maker declares. And the count is 0, where How to set up Ollama locally and see what it sends out finds 30 calls a day with the app open.

The order that leaves you with a working and quiet MLX:

  1. Check that your Python is native and 3.11 or newer: platform.processor() has to say arm.
  2. Create a virtual environment, pip install mlx-lm, and confirm 0.32.0 and 0.32.3 with pip show mlx mlx-lm.
  3. Pick the model with the model finder or your Mac’s hardware sheet, and read Peak memory in mlx_lm.generate.
  4. Once it is downloaded, set HF_HUB_OFFLINE=1 and HF_HUB_DISABLE_TELEMETRY=1.
  5. Before running mlx_lm.server, set --allowed-origins, leave the host at 127.0.0.1, and use an SSH tunnel for other machines.

The 5 runtimes are compared on exactly this in 5 local AI runtimes: telemetry, network calls and audit.