12. Advanced (1): LLM Providers & Models

omicOS's agents are powered by large language models under the hood. The models provided by the omicOS cloud work fine by default; this chapter covers advanced usage — bringing your own API key and pinning a specific provider and model.

12.1 Provider Selection Priority

When omicos decides "which vendor to use," it tries the following in order and stops at the first match:

  1. The provider explicitly specified in the chat request (chosen in the web app's settings, or passed to the API)
  2. The OMICOS_LLM_PROVIDER environment variable
  3. The OMICOS_PROVIDER environment variable (legacy alias)
  4. Inferred from the model name (e.g. a model name containing deepseek uses deepseek)
  5. Inferred from the first API key found, in this fixed order: DEEPSEEK_API_KEY → MINIMAX_API_KEY → OPENAI_API_KEY

If none of these match, it errors:

Error: no model provider configured

The mock provider is explicitly disabled — specifying it (or a model name with the mock/ prefix) errors immediately:

Error: mock provider is disabled; configure a real model provider

12.2 Model Selection Priority

  1. The model explicitly specified in the chat request
  2. OMICOS_LLM_MODEL
  3. OMICOS_MODEL (legacy alias)
  4. The provider's default: deepseek → deepseek-v4-flash, everything else → gpt-4o-mini

The model name may carry a provider/ prefix, which is stripped automatically — e.g. deepseek/deepseek-v4-flash is equivalent to deepseek-v4-flash.

12.3 Supported Providers

Category Examples Notes
OpenAI-compatible (catalog-driven) openai, deepseek, qwen, zhipu, moonshot, xai, groq, mistral, ollama, openrouter, together, fireworks, deepinfra, cerebras, perplexity, minimax, siliconflow, and more The most common case, supplied dynamically by the cloud model catalog
Anthropic anthropic Uses the Anthropic Messages protocol; the key comes from ANTHROPIC_API_KEY, and OAuth login is also supported
OAuth codex (OpenAI), gemini-cli (Google), anthropic-oauth, xai-oauth Uses a third-party account's OAuth credentials, no API key needed
Custom custom_openai, custom_anthropic Points at your own self-hosted / private endpoint

omicos does not send a temperature parameter to the server — some reasoning models reject temperature ≠ 1, so it's simply never sent.

12.4 Where the API Key Comes From

When resolving the API key for a given provider, omicos looks in this order:

  1. The environment variable <PROVIDER>_API_KEY, with hyphens in the provider id converted to underscores. For example, alibaba-coding-plan maps to ALIBABA_CODING_PLAN_API_KEY.
  2. The auth.json file, searched in this order for the first one that exists and has the field:
    • $OMICOS_LOCAL_HOME/auth.json
    • $OMICOS_RUNTIME_HOME/auth.json
    • ~/.omicos/auth.json
    • <current directory>/.omicos/auth.json (a fallback kept only for backward compatibility)

auth.json is a simple JSON dictionary; only non-empty values take effect:

{
  "DEEPSEEK_API_KEY": "sk-...",
  "OPENAI_API_KEY": "sk-..."
}

Ollama is a special case: when no key is set it automatically uses the placeholder ollama, since a local Ollama instance doesn't need a real key.

Once configured, confirm it with a quick run (swap in your own provider / model):

export DEEPSEEK_API_KEY=sk-your-key
export OMICOS_LLM_PROVIDER=deepseek
omicos serve --no-browser 2>&1 | head -3

Expected output:

[omicos] online: my-analysis (local-9f2c4d7a1b8e4f5c8d3a6b0e2f7c1a94)
[omicos] listening: http://127.0.0.1:5055

A misconfigured key doesn't fail startup — the provider is only resolved when you send your first message, and any error shows up there in the chat.

12.5 Custom Endpoints

Each provider's endpoint is resolved in this order:

  1. The <PROVIDER>_API_BASE environment variable
  2. api_base from the cloud model catalog (cached at ~/.omicos/cloud-models/models.json)

Special case: custom_openai's default endpoint is http://127.0.0.1:8000/v1, convenient for connecting to a local inference service (like vLLM). Its key comes from CUSTOM_OPENAI_API_KEY, falling back to OPENAI_API_KEY if unset.

A complete example — connecting to a self-hosted vLLM service:

export OMICOS_LLM_PROVIDER=custom_openai
export CUSTOM_OPENAI_API_BASE=http://127.0.0.1:8000/v1
export CUSTOM_OPENAI_API_KEY=dummy           # local services usually don't validate the key
export OMICOS_LLM_MODEL=Qwen2.5-72B-Instruct
omicos serve --no-browser 2>&1 | head -3

Expected output:

[omicos] online: my-analysis (local-9f2c4d7a1b8e4f5c8d3a6b0e2f7c1a94)
[omicos] listening: http://127.0.0.1:5055

custom_anthropic works the same way, using CUSTOM_ANTHROPIC_API_BASE for the endpoint and CUSTOM_ANTHROPIC_API_KEY (falling back to ANTHROPIC_API_KEY) for the key. It has no default endpoint — leaving it unset errors immediately, so you know what's missing.

12.6 The Cloud Model Catalog

omicos fetches a model catalog from the cloud (which models are available, their context windows, whether they support vision, etc.), cached at ~/.omicos/cloud-models/models.json. Related variables:

Variable Purpose
OMICOS_MODELS_OFFLINE Offline mode, use only the local cache
OMICOS_MODELS_CLOUD_URL Override the catalog fetch address
OMICOS_MODELS_CACHE_DIR Override the cache directory

To check what's actually in the cache:

python3 -c "import json;d=json.load(open('$HOME/.omicos/cloud-models/models.json'));print(len(d['providers']),'providers; first:',d['providers'][0]['id'])"

Expected output (the provider count depends on the cloud catalog's contents at the time):

23 providers; first: omicos-cloud

The omicos-cloud entry ranked first is the built-in cloud provider — it's always placed at the front regardless of what else is in the catalog.

12.7 Vision Models

If you need a model that can "look at" images (for example, to interpret a generated chart), you can configure a separate vision model, independent of the main chat model:

export OMICOS_VISION_MODEL=gpt-4o
export OMICOS_VISION_BASE_URL=https://api.openai.com/v1
export OMICOS_VISION_API_KEY=sk-...
omicos serve --no-browser 2>&1 | head -2

Expected output:

[omicos] online: my-analysis (local-9f2c4d7a1b8e4f5c8d3a6b0e2f7c1a94)
[omicos] listening: http://127.0.0.1:5055

results matching ""

    No results matching ""