Skip to content

Custom endpoint

Olyvr works with any service that offers the OpenAI chat completions API with streaming and tool calling. That includes most hosted providers, AI gateways, and self-hosted servers such as llama.cpp, vLLM or Ollama’s OpenAI-compatible mode.

  1. Open Settings → Model access and expand Endpoint & advanced.

  2. Enter the base URL, ending in /v1 for most providers, for example https://api.example.com/v1.

  3. Enter the API key if the endpoint needs one, and save.

  4. In a chat, open the composer’s model picker. Olyvr lists the models the endpoint reports at <base URL>/models.

The key is encrypted on disk and is never shown again after you save it. To change it, paste a new key; to remove it, clear the key.

For managed machines or scripted setups, set these before starting Olyvr. They take priority over the values saved in Settings and are never written to disk:

Variable Example
ATLAS_MODEL_BASE https://api.example.com/v1
ATLAS_MODEL_KEY sk-…
ATLAS_MODEL gpt-5.5

You pick the model per conversation in the composer. When there is no choice yet, Olyvr prefers a model whose name starts with gpt- or claude-. Endpoints on your own machine (localhost or 127.0.0.1) are labelled as local.

Message Meaning
“Configure a model endpoint in Settings → Model access…” No endpoint is set. Add one, or use a subscription.
“Model endpoint returned HTTP 401/403…” The API key is wrong or lacks access to that model.
“Model endpoint returned HTTP 400/404…” The model name is wrong, or the endpoint doesn’t support streaming or tool calls.
“The model returned no answer or tool calls.” The model replied with nothing usable. Try again or choose a more capable model.

Requests time out after two minutes.