12. Advanced (1): LLM Providers & Models
omicOS's agents are powered by large language models under the hood. The models provided by the omicOS cloud work fine by default; this chapter covers advanced usage — bringing your own API key and pinning a specific provider and model.
12.1 Provider Selection Priority
When omicos decides "which vendor to use," it tries the following in order and stops at the first match:
- The provider explicitly specified in the chat request (chosen in the web app's settings, or passed to the API)
- The
OMICOS_LLM_PROVIDERenvironment variable - The
OMICOS_PROVIDERenvironment variable (legacy alias) - Inferred from the model name (e.g. a model name containing
deepseekuses deepseek) - Inferred from the first API key found, in this fixed order:
DEEPSEEK_API_KEY→MINIMAX_API_KEY→OPENAI_API_KEY
If none of these match, it errors:
Error: no model provider configured
The mock provider is explicitly disabled — specifying it (or a model name with the mock/ prefix) errors immediately:
Error: mock provider is disabled; configure a real model provider
12.2 Model Selection Priority
- The model explicitly specified in the chat request
OMICOS_LLM_MODELOMICOS_MODEL(legacy alias)- The provider's default: deepseek →
deepseek-v4-flash, everything else →gpt-4o-mini
The model name may carry a provider/ prefix, which is stripped automatically — e.g. deepseek/deepseek-v4-flash is equivalent to deepseek-v4-flash.
12.3 Supported Providers
| Category | Examples | Notes |
|---|---|---|
| OpenAI-compatible (catalog-driven) | openai, deepseek, qwen, zhipu, moonshot, xai, groq, mistral, ollama, openrouter, together, fireworks, deepinfra, cerebras, perplexity, minimax, siliconflow, and more | The most common case, supplied dynamically by the cloud model catalog |
| Anthropic | anthropic | Uses the Anthropic Messages protocol; the key comes from ANTHROPIC_API_KEY, and OAuth login is also supported |
| OAuth | codex (OpenAI), gemini-cli (Google), anthropic-oauth, xai-oauth | Uses a third-party account's OAuth credentials, no API key needed |
| Custom | custom_openai, custom_anthropic | Points at your own self-hosted / private endpoint |
omicos does not send a
temperatureparameter to the server — some reasoning models rejecttemperature ≠ 1, so it's simply never sent.
12.4 Where the API Key Comes From
When resolving the API key for a given provider, omicos looks in this order:
- The environment variable
<PROVIDER>_API_KEY, with hyphens in the provider id converted to underscores. For example,alibaba-coding-planmaps toALIBABA_CODING_PLAN_API_KEY. - The
auth.jsonfile, searched in this order for the first one that exists and has the field:$OMICOS_LOCAL_HOME/auth.json$OMICOS_RUNTIME_HOME/auth.json~/.omicos/auth.json<current directory>/.omicos/auth.json(a fallback kept only for backward compatibility)
auth.json is a simple JSON dictionary; only non-empty values take effect:
{
"DEEPSEEK_API_KEY": "sk-...",
"OPENAI_API_KEY": "sk-..."
}
Ollama is a special case: when no key is set it automatically uses the placeholder
ollama, since a local Ollama instance doesn't need a real key.
Once configured, confirm it with a quick run (swap in your own provider / model):
export DEEPSEEK_API_KEY=sk-your-key
export OMICOS_LLM_PROVIDER=deepseek
omicos serve --no-browser 2>&1 | head -3
Expected output:
[omicos] online: my-analysis (local-9f2c4d7a1b8e4f5c8d3a6b0e2f7c1a94)
[omicos] listening: http://127.0.0.1:5055
A misconfigured key doesn't fail startup — the provider is only resolved when you send your first message, and any error shows up there in the chat.
12.5 Custom Endpoints
Each provider's endpoint is resolved in this order:
- The
<PROVIDER>_API_BASEenvironment variable api_basefrom the cloud model catalog (cached at~/.omicos/cloud-models/models.json)
Special case: custom_openai's default endpoint is http://127.0.0.1:8000/v1, convenient for connecting to a local inference service (like vLLM). Its key comes from CUSTOM_OPENAI_API_KEY, falling back to OPENAI_API_KEY if unset.
A complete example — connecting to a self-hosted vLLM service:
export OMICOS_LLM_PROVIDER=custom_openai
export CUSTOM_OPENAI_API_BASE=http://127.0.0.1:8000/v1
export CUSTOM_OPENAI_API_KEY=dummy # local services usually don't validate the key
export OMICOS_LLM_MODEL=Qwen2.5-72B-Instruct
omicos serve --no-browser 2>&1 | head -3
Expected output:
[omicos] online: my-analysis (local-9f2c4d7a1b8e4f5c8d3a6b0e2f7c1a94)
[omicos] listening: http://127.0.0.1:5055
custom_anthropic works the same way, using CUSTOM_ANTHROPIC_API_BASE for the endpoint and CUSTOM_ANTHROPIC_API_KEY (falling back to ANTHROPIC_API_KEY) for the key. It has no default endpoint — leaving it unset errors immediately, so you know what's missing.
12.6 The Cloud Model Catalog
omicos fetches a model catalog from the cloud (which models are available, their context windows, whether they support vision, etc.), cached at ~/.omicos/cloud-models/models.json. Related variables:
| Variable | Purpose |
|---|---|
OMICOS_MODELS_OFFLINE |
Offline mode, use only the local cache |
OMICOS_MODELS_CLOUD_URL |
Override the catalog fetch address |
OMICOS_MODELS_CACHE_DIR |
Override the cache directory |
To check what's actually in the cache:
python3 -c "import json;d=json.load(open('$HOME/.omicos/cloud-models/models.json'));print(len(d['providers']),'providers; first:',d['providers'][0]['id'])"
Expected output (the provider count depends on the cloud catalog's contents at the time):
23 providers; first: omicos-cloud
The omicos-cloud entry ranked first is the built-in cloud provider — it's always placed at the front regardless of what else is in the catalog.
12.7 Vision Models
If you need a model that can "look at" images (for example, to interpret a generated chart), you can configure a separate vision model, independent of the main chat model:
export OMICOS_VISION_MODEL=gpt-4o
export OMICOS_VISION_BASE_URL=https://api.openai.com/v1
export OMICOS_VISION_API_KEY=sk-...
omicos serve --no-browser 2>&1 | head -2
Expected output:
[omicos] online: my-analysis (local-9f2c4d7a1b8e4f5c8d3a6b0e2f7c1a94)
[omicos] listening: http://127.0.0.1:5055