Home›Guides›How to set up Ollama locally and see what it sends out

Setup · Guide

How to set up Ollama locally and see what it sends out

Install Ollama, let the app or ollama serve start the server, and open http://127.0.0.1:11434/: if it answers “Ollama is running”, it works, and only your machine can reach it. Quiet it is not. With the desktop app open, Ollama calls ollama.com 30 times a day on its own, and 24 of them cannot be switched off.

Measured and written by Local AI ScopePublished Figures measured by Local AI Scope

Install Ollama, let the app or ollama serve start the server, and open http://127.0.0.1:11434/: if it answers “Ollama is running”, it works, and only your machine can reach it. Quiet it is not. With the desktop app open, Ollama calls ollama.com 30 times a day on its own, and 24 of them cannot be switched off.

The install takes a few minutes. The decisions that matter come after it: which address the server listens on, what it sends out while you are not using it, and where your models and chats end up. This page goes through the setup in that order, with the commands of the current stable version, 0.35.0 (published 28 September 2026). We read them on 0.34.4, and none of them changes in 0.35.0: its code differs in the desktop app, the benchmark tool and a new endpoint for decision models, not in the commands, variables or install script used here.

How to install Ollama on Mac, Windows and Linux

There is one line per system, and each one leaves you with something slightly different.

System Install command What you end up with
macOS 14 Sonoma or newer curl -fsSL https://ollama.com/install.sh | sh or the Ollama.dmg installer Desktop app in Applications, plus the ollama command. GPU on Apple Silicon, CPU only on Intel Macs
Windows 10 22H2 or newer irm https://ollama.com/install.ps1 | iex in PowerShell, or OllamaSetup.exe Desktop app in your user folder, no administrator rights needed, ollama in any terminal
Linux curl -fsSL https://ollama.com/install.sh | sh A systemd service called ollama that runs as its own ollama user. No desktop app
Docker docker run -d -v ollama:/root/.ollama -p 127.0.0.1:11434:11434 --name ollama ollama/ollama The server in a container. Note the 127.0.0.1: the reason is further down

Commands and requirements as published by Ollama in its README and its macOS, Windows and Linux pages for version 0.34.4, read on 25 September 2026. The Docker line is Ollama’s own, with the port bound to 127.0.0.1.

Read the script before you pipe it. https://ollama.com/install.sh redirects to the install.sh attached to the latest release on GitHub; on 25 September 2026 it was byte for byte the one tagged v0.34.4. On a Mac the same script installs the desktop app, which matters for the next sections: the app is what adds the hourly update check.

ollama serve: do you need to run it yourself?

Usually not. On macOS and Windows the desktop app launches ollama serve for you when it starts. On Linux the installer creates a service whose only job is ollama serve, and sudo systemctl enable ollama keeps it running across reboots.

You type it by hand in three cases: a manual install from the tarball, a server you want to watch live in the terminal, or a one-off run with different settings:

OLLAMA_DEBUG=1 ollama serve

ollama serve --help lists the environment variables the server reads. If the app or the service is already running, a second ollama serve fights it for port 11434. Quit one first.

How to check if Ollama is running

Four checks, from the most basic to the most useful.

The server answers. Open http://127.0.0.1:11434/ in a browser, or from a terminal:

curl http://127.0.0.1:11434/
Ollama is running

Which version is answering. The API has an endpoint for it, and the command line asks the same question:

curl http://127.0.0.1:11434/api/version
{"version":"0.35.0"}

ollama -v
ollama version is 0.35.0

If ollama -v prints Warning: could not connect to a running Ollama instance, the command is installed but no server is up. If it adds Warning: client version is …, the command and the server are different versions, which usually means two installs.

What is loaded, and where. ollama ps lists the models in memory with the columns NAME, ID, SIZE, PROCESSOR, CONTEXT and UNTIL. The PROCESSOR column is the one to read: 100% GPU means the whole model is on the graphics card, 100% CPU means it is in system memory, and a split such as 48%/52% CPU/GPU means it did not fit and is running slower than it could.

The service, on Linux. sudo systemctl status ollama tells you whether the service is active and since when.

Verbose mode and logs: two different things

The verbose flag is about speed. ollama run <model> --verbose, or /set verbose inside a chat, prints timings after every answer: total duration, load duration, prompt eval rate and eval rate, the last two in tokens per second. /set quiet turns it off again.

Debug logging is about the server. OLLAMA_DEBUG=1 raises the log to debug level and OLLAMA_DEBUG=2 to trace. How you set it depends on how Ollama runs:

# Linux service: add Environment="OLLAMA_DEBUG=1" under [Service]
sudo systemctl edit ollama
sudo systemctl daemon-reload
sudo systemctl restart ollama

# macOS app: set it, then quit and reopen Ollama
launchctl setenv OLLAMA_DEBUG 1

# Windows: quit the app from the tray first, then in PowerShell
$env:OLLAMA_DEBUG="1"
& "ollama app.exe"

The logs themselves live in ~/.ollama/logs/server.log on a Mac, in %LOCALAPPDATA%\Ollama\server.log on Windows, in journalctl -u ollama --no-pager --follow --pager-end on Linux, and in docker logs ollama in a container.

Flags, variables and log paths from Ollama’s troubleshooting and FAQ pages and from its source code, version 0.34.4, read on 25 September 2026.

Which address Ollama listens on, and why it has no password

127.0.0.1:11434 out of the box. Only programs on your own machine can reach it. That is the right default, and the one worth keeping.

OLLAMA_HOST=0.0.0.0:11434 makes it listen on every network interface, and there is a catch that the variable’s name does not suggest: the local API has no authentication of any kind. OLLAMA_AUTH does not change that; it only makes the command-line client sign its own requests. And once the server is not on the loopback address, it also stops checking the Host header, which is one of its two remaining defences. Anyone who can reach that port can use your models and your graphics card. The desktop app has a setting that does exactly this, sets OLLAMA_HOST=0.0.0.0, and it comes switched off.

If you need it from another machine, leave the server where it is and carry the port over SSH instead:

ssh -L 11434:127.0.0.1:11434 you@the-machine-running-ollama

While that session is open, http://127.0.0.1:11434/ on your laptop is the remote Ollama, and the remote Ollama is still closed to everyone else.

Docker is the easy place to get this wrong. Ollama’s image sets OLLAMA_HOST=0.0.0.0:11434 inside the container, which it needs, and a plain -p 11434:11434 publishes the port on all of the host’s addresses, according to Docker’s own documentation. Hence the 127.0.0.1: in front of the port in the table above.

The other defence is CORS, and it is looser than the address suggests: without any configuration, Ollama accepts browser requests from pages served on localhost on any port, from files opened in the browser (file://) and from browser extensions. OLLAMA_ORIGINS adds to that list; it does not replace it.

What Ollama sends out on its own, and what you can switch off

This is the part the install screen does not mention. With the program open and you doing nothing, these are the calls it makes:

Call Where to Per day Can you switch it off?
Update check, signed with a key stored on your machine; on macOS it also carries a persistent device identifier ollama.com/api/update 24 (every hour, while the desktop app is open; on Linux there is none) No. The auto-update setting only stops the download, not the check
Remote catalogue of recommended models, started by ollama serve itself: no desktop app and no account needed ollama.com/api/experimental/model-recommendations 6 (every 4 hours) Yes: OLLAMA_NO_CLOUD=1

Counted by Local AI Scope from the cadence written in Ollama’s published source code (version 0.32.15, read on 25 August 2026), not from a traffic capture. The two intervals are the same in the code of 0.34.4, which we read again on 25 September 2026. Full sheet with every source: Ollama: what it sends home and whether you can audit it.

That adds up to 30 calls a day on a Mac or a Windows PC with the app open, 24 of which you cannot prevent. A Linux server with no desktop app makes 6, and all 6 can be switched off. Downloading a model is not counted here: that happens when you ask for it, from registry.ollama.ai, and for public models it goes out without any credential.

The catalogue request is signed. Since version 0.34.4, the 4-hour call to the catalogue goes out signed with ~/.ollama/id_ed25519, the key that identifies your installation to ollama.com and the same one that signs the update check. That happens on a Linux server with no desktop app too, and OLLAMA_NO_CLOUD=1, the switch that comes next, cuts it. What it does not cut is the account check the desktop app makes when it opens: a signed request to ollama.com/api/me, even if you have no account. It is not among the 30 because it runs on no clock: it goes out when you open the app.

How to switch off what can be switched off. On the Linux service, add Environment="OLLAMA_NO_CLOUD=1" the same way as the debug line above. On a Mac or a Windows PC, write this file, which both the server and the desktop app read:

# ~/.ollama/server.json
{
  "disable_ollama_cloud": true
}

Restart Ollama and look for Ollama cloud disabled: true in the log. That line is your confirmation. The same switch turns off the :cloud models and web search, which is the point: a model whose name ends in :cloud does not run on your machine. It is the same command line and the same catalogue, and your prompt goes to Ollama’s servers. If you only use local models, you never need ollama signin either. And if you would rather the command line kept no record of what you type, OLLAMA_NOHISTORY=1 stops it writing ~/.ollama/history.

None of these calls carries your prompts. Ollama’s privacy policy (March 2026) says it does not see your prompts or data when you run locally, and that it may collect limited device and usage metadata, such as the app version and request counts. That is the vendor’s word, and the code we read is consistent with it. What we have not done is put a network analyser in front of the program: this is a reading of the code, not a recording of the wire.

Where Ollama keeps your models and chats

Models take tens to hundreds of gigabytes, so the location matters before the first download.

What Where
Models, macOS ~/.ollama/models
Models, Windows C:\Users\%username%\.ollama\models
Models, Linux service /usr/share/ollama/.ollama/models
Key that identifies your install to ollama.com ~/.ollama/id_ed25519 and .pub (Linux service: /usr/share/ollama/.ollama/)
Desktop app chats, an unencrypted SQLite database macOS ~/Library/Application Support/Ollama/db.sqlite · Windows %LOCALAPPDATA%\Ollama\db.sqlite

Paths as published in Ollama’s FAQ and read in its source code, version 0.34.4, on 25 September 2026.

OLLAMA_MODELS moves the models anywhere you like. On Linux, the ollama user needs to own the new folder: sudo chown -R ollama:ollama <folder>. The chat database never leaves the machine, and anyone who can open files in your user account can read it.

What to download first

Pick the model after the setup, not before. The size of the file has to fit in the memory you have left, together with the conversation, and a model that spills over to the processor shows up as that split in ollama ps. The model finder tells you which models fit your machine at a usable compression and how long a document they can take, and it runs entirely in your browser. For what a local model can and cannot do compared with the one you use today, What AI your computer can run, and what you give up to run it has the numbers.

The order that leaves you with a working and quiet Ollama:

  1. Install it with the line for your system.
  2. Check http://127.0.0.1:11434/ and ollama -v.
  3. Write ~/.ollama/server.json with disable_ollama_cloud (on the Linux service, the Environment line instead), restart, and look for the confirmation line in the log.
  4. Leave OLLAMA_HOST alone; use an SSH tunnel if another machine needs it.
  5. Pull a model that fits, run it with --verbose, and check ollama ps says 100% GPU.

If the hourly update check is the part you cannot live with, it belongs to the desktop app and not to the server. The five runtimes are compared on exactly this in 5 local AI runtimes: telemetry, network calls and audit: llama.cpp, for one, makes no call of its own, and LM Studio is the only one of the five whose code no one outside can check.