Skip to content

Local models

Olyvr can download an open-weight model and run it on your computer with llama.cpp. Nothing is sent to a cloud provider.

Model Variant Download Recommended RAM
Qwen3.5 9B Q4 (UD-Q4_K_XL) 6.1 GB 16 GB
Qwen3.5 9B Q8_0 9.8 GB 24 GB
gpt-oss 20B F16 13.7 GB 24 GB
Qwen3.6 35B-A3B UD-IQ3_XXS 14.1 GB 24 GB
Qwen3.6 35B-A3B UD-Q4_K_M 22.7 GB 32 GB

If you don’t choose a variant, Olyvr picks one that fits your machine’s memory. The Qwen models can also read images.

  1. Open Settings → Model access → Local models.

  2. Choose a model and click install. Olyvr downloads the llama.cpp runtime from its GitHub releases and the model from Hugging Face. Interrupted downloads resume where they left off.

  3. When the download finishes, the model server starts automatically. Starting can take a few minutes for large models.

Models and the runtime are stored in your data folder, under state/local. Model files are large, so check your free disk space before installing.

If you already use Hermes to manage llama.cpp models, Olyvr detects a running Hermes llama.cpp server and its downloaded models (in ~/.hermes, or $HERMES_HOME) and can reuse them instead of downloading again.

These environment variables are read when the model server starts:

Variable Default Effect
ATLAS_LOCAL_CTX 131072 Context length in tokens
ATLAS_LOCAL_SLOTS 2 Parallel request slots
ATLAS_LOCAL_VISION off Set to 1 to load the vision projector
ATLAS_LOCAL_EXTRA_FLAGS none Extra llama-server arguments
  • macOS: the downloaded runtime is built for Apple Silicon.
  • Windows and Linux: local models have had less testing than on macOS. If the server does not start, check server.log in state/local and report the problem.

If startup takes longer than 4 minutes, Olyvr stops waiting and points you to the log file.