# Use local AI models

> Run every model Oatmilk uses on your own hardware with Ollama, LM Studio, vLLM or llama.cpp.

Source: https://app.getoatmilk.com/docs/self-hosting/local-models

Oatmilk uses AI models to read statements, receipts and documents, to sort mail and match receipts, and for Ask AI. By default it calls hosted models through Vercel AI Gateway. Pointed at a model server on your own hardware instead, every one of those calls stays with you, nothing is billed, and no API key is needed.

This works with [Ollama](https://ollama.com), [LM Studio](https://lmstudio.ai) and any server that speaks the OpenAI chat completions API, such as vLLM, llama.cpp's `llama-server` and LocalAI.

## What your computer needs

A rough guide to the memory the models need, on top of Oatmilk's own 4 GB:

| Setup | Memory | Good for |
| --- | --- | --- |
| A 9B chat model and a 4B classifier | 16 GB, more is better | Real books on one computer |
| 4B models | 8 GB | Trying it out; slower and less accurate |
| 0.8B models | 4 GB | Checking that everything is connected, not real use |
| 30B models and larger, on a GPU server | 32 GB of GPU memory or more | A team, faster and more accurate |

A Mac with Apple silicon, or an NVIDIA GPU, makes answers much faster than a CPU alone.

## Option 1: Ollama

Use Ollama 0.35 or newer, which adds decision models that answer Oatmilk's classifier questions with a confidence, as the hosted classifier does.

### 1. Install Ollama and pull the models

```bash
ollama pull qwen3.5:9b   # chat, tools, extraction and receipt images
ollama pull tev1         # classifiers: a decision model (nimble is larger and more accurate)
```

### 2. Start it with a longer context

Ollama loads models with a short context by default. Long statements and the agents need more:

```bash
OLLAMA_CONTEXT_LENGTH=32768 ollama serve
```

On Linux, Oatmilk's container reaches your computer through Docker's network, not `127.0.0.1`, so Ollama must listen on every address:

```bash
OLLAMA_HOST=0.0.0.0 OLLAMA_CONTEXT_LENGTH=32768 ollama serve
```

With Ollama's systemd service, set both with `sudo systemctl edit ollama` instead. Docker Desktop on macOS and Windows reaches Ollama either way.

### 3. Point Oatmilk at it

In `bun run self-host setup`, choose **Ollama on this computer**, and keep `qwen3.5:9b` and `tev1` as the models. Setup writes:

```bash title="self-host/.env"
OATMILK_MODEL_PROVIDER=ollama
OATMILK_LOCAL_BASE_URL=http://host.docker.internal:11434/v1
OATMILK_LOCAL_MODEL=qwen3.5:9b
OATMILK_LOCAL_CLASSIFIER_MODEL=tev1
OATMILK_LOCAL_CLASSIFIER_API=systemone
```

On an older Ollama, remove the last line, and the chat model answers the classifier questions too.

## Option 2: LM Studio

1. Download a model in LM Studio, for example **Qwen3.5 9B** (it reads images too) and, for classifiers, **Qwen3.5 4B**. Set each one's context length to 32768 or more when you load it.
2. Open the **Developer** tab and start the server, or run `lms server start`. It listens on port 1234.
3. On Linux, turn on **Serve on Local Network** in the server settings, so Oatmilk's container can reach it.
4. In `bun run self-host setup`, choose **LM Studio on this computer**, with the model identifiers LM Studio shows.

```bash title="self-host/.env"
OATMILK_MODEL_PROVIDER=lmstudio
OATMILK_LOCAL_BASE_URL=http://host.docker.internal:1234/v1
OATMILK_LOCAL_MODEL=qwen/qwen3.5-9b
OATMILK_LOCAL_CLASSIFIER_MODEL=qwen/qwen3.5-4b
OATMILK_LOCAL_CLASSIFIER_API=chat
```

## Option 3: vLLM, llama.cpp or another server

Any server with an OpenAI-compatible `/v1` address works, on this computer or another one on your network. For example, vLLM on a GPU server:

```bash
vllm serve Qwen/Qwen3-8B --enable-auto-tool-choice --tool-call-parser hermes --max-model-len 32768
```

In setup, choose **Another OpenAI-compatible server** and give its address, ending in `/v1`, and the model's name. An address typed as `localhost` is pointed back at your computer from inside the container.

```bash title="self-host/.env"
OATMILK_MODEL_PROVIDER=openai-compatible
OATMILK_LOCAL_BASE_URL=http://192.168.1.50:8000/v1
OATMILK_LOCAL_MODEL=Qwen/Qwen3-8B
OATMILK_LOCAL_API_KEY=            # only if your server asks for one
```

Choose a model that supports tool calling and structured output, and start the server with tool calling turned on.

## Change models later

Edit `self-host/.env`, then apply it:

```bash
bun run self-host up
bun run self-host doctor   # checks that the app container reaches the model server
```

## Settings

| Setting | Default | What it does |
| --- | --- | --- |
| `OATMILK_MODEL_PROVIDER` | `gateway` | `gateway`, `ollama`, `lmstudio` or `openai-compatible` |
| `OATMILK_LOCAL_BASE_URL` | Ollama's or LM Studio's usual address | The server's `/v1` address. Required for `openai-compatible`. |
| `OATMILK_LOCAL_API_KEY` | none | Sent as a bearer token when set |
| `OATMILK_LOCAL_MODEL` | required | Answers every chat, tool, extraction and review call |
| `OATMILK_LOCAL_VISION_MODEL` | the chat model | Used whenever a prompt has an image, a PDF page or another file |
| `OATMILK_LOCAL_REASONING` | each call's own | Caps thinking: `none`, `minimal`, `low`, `medium`, `high` or `xhigh`. `none` answers fastest on a laptop. |
| `OATMILK_LOCAL_CLASSIFIER_MODEL` | the chat model | Answers classifier questions: document kinds, mail routing, receipt matches, categories |
| `OATMILK_LOCAL_CLASSIFIER_API` | `chat` | `systemone` for an Ollama decision model, `chat` to ask a chat model |
| `OATMILK_LOCAL_CLASSIFIER_BASE_URL` | `OATMILK_LOCAL_BASE_URL` | When classifiers run on another server, such as Ollama next to LM Studio |
| `OATMILK_LOCAL_EMBEDDING_MODEL` | none | For embeddings, for example `nomic-embed-text` |
| `OATMILK_LOCAL_TRANSCRIPTION_MODEL` | none | Voice input in Ask AI, from a server with `/audio/transcriptions` such as a whisper.cpp server |
| `OATMILK_LOCAL_TRANSCRIPTION_BASE_URL` | `OATMILK_LOCAL_BASE_URL` | When transcription runs on another server |
| `OATMILK_LOCAL_MODEL_MAP` | none | Pins particular hosted model ids to particular local models, such as `typesafe-ai/jev=qwen3:4b` |

An invalid setting stops Oatmilk at start-up with a message that names the setting to fix.

## Classifiers and automatic matching

Oatmilk's classifiers answer typed questions with probabilities. An Ollama decision model (`systemone`) also reports a confidence for each answer, the way the hosted classifier does, so automatic steps such as matching a receipt to a bank line keep working: they still need 98% probability and 95% confidence.

A chat model (`chat`) gives its own estimate of the probabilities but no confidence. Every classifier still works, but steps that need confidence, such as automatic receipt matching, leave the decision to a person.

## Which models to choose

| Role | Ollama | LM Studio |
| --- | --- | --- |
| Chat, tools and extraction | `qwen3.5:9b` (smaller: `qwen3.5:4b`, `gemma4:e4b`) | `qwen/qwen3.5-9b`, `google/gemma-4-e4b` |
| Classifiers | `nimble` or `tev1`, with `systemone` | `qwen/qwen3.5-4b`, with `chat` |
| Receipts and documents (vision) | `qwen3.5:9b`, `gemma4:12b` | `qwen/qwen3.5-9b`, `google/gemma-4-12b` |
| Embeddings | `nomic-embed-text` | nomic-embed-text v1.5 |

Qwen3.5 and similar models think before they answer. On slow hardware, set `OATMILK_LOCAL_REASONING=none`.

## Limits

- Quality depends on the model. Check a few real statements and receipts before you trust a model with your books.
- Image generation isn't available locally, and voice input needs a transcription server.
- Some agents look up their model's context window in a public catalog. If you run fully offline and an agent doesn't start, pin its model to a local one with `OATMILK_LOCAL_MODEL_MAP` and check `bun run self-host logs app`.

## Without the Docker stack

The same settings work for `bun run dev` in `.env.local`. There, Oatmilk runs on your computer itself, so use `http://localhost:11434/v1` (Ollama) or `http://localhost:1234/v1` (LM Studio), which are also the defaults. See [Develop Oatmilk locally](https://app.getoatmilk.com/docs/self-hosting/develop.md).
