Skip to main content

Provider and Model Configuration

Every field for choosing and reaching a model. For which provider to pick and why, read Providers and Models and the Provider Cookbook.

This page is one of four. The configuration reference was a single 2,929-line document, and its own second section was a hand-written table of contents. That is a page admitting it could not be navigated. It is now split by subject, so the fields for a thing sit beside the pages that explain the thing:

PageHolds
Configurationhow config is loaded, secrets, environment variables, channels, and the deployment-wide settings
Provider and Model Configurationevery provider, model preset, fallback and transcription field
Agent and Tool Configurationnamed agents, tool groups, web tools, MCP, knowledge and connectors
Security Configurationcapability gates, approvers, standing grants and pairing

Examples are snippets to merge into ~/.nanoinfra/config.json, not replacement files. The docs use camelCase because nanoinfra writes config that way.

Providers​

[!TIP]

  • Voice transcription: Voice messages and WebUI microphone input use the shared top-level transcription settings. The default transcription.provider value is "groq". Set it to "openai" for OpenAI Whisper, "openrouter" for OpenRouter speech-to-text models, "xiaomi_mimo" for Xiaomi MiMo ASR, or "assemblyai" for AssemblyAI. API keys still live in the matching providers.<provider> config.
  • MiniMax Coding Plan: Exclusive discount links for the nanoinfra community: Overseas · Mainland China.
  • MiniMax (Mainland China): If your API key is from MiniMax's mainland China platform (minimaxi.com), set "apiBase": "https://api.minimaxi.com/v1" in your minimax provider config.
  • MiniMax thinking mode: providers.minimaxAnthropic is the config block for reasoningEffort / thinking mode. MiniMax exposes that capability through its Anthropic-compatible endpoint. So nanoinfra keeps it as a separate provider, rather than guessing MiniMax-specific thinking parameters on the generic OpenAI-compatible minimax endpoint. It uses the same MINIMAX_API_KEY. Default Anthropic-compatible base URL: https://api.minimax.io/anthropic. For mainland China use https://api.minimaxi.com/anthropic.
  • Kimi Coding Plan: Use providers.kimiCoding with provider: "kimi_coding" for Kimi's dedicated Anthropic Messages API endpoint. The endpoint requires a Claude-compatible User-Agent. nanoinfra sends claude-code/0.1.0 by default, and you can override it with extraHeaders.User-Agent if your account requires a different value.
  • VolcEngine / BytePlus Coding Plan: Subscription endpoints are configured through dedicated providers volcengineCodingPlan or byteplusCodingPlan, separate from the pay-per-use volcengine / byteplus providers.
  • OpenCode Zen / Go: providers.opencode (canonical Zen), the legacy-compatible providers.opencodeZen, and providers.opencodeGo use the same OPENCODE_API_KEY, but route to different OpenCode gateways. These providers use OpenCode's OpenAI-compatible chat/completions endpoints. Choose model IDs from that endpoint family.
  • Zhipu Coding Plan: If you're on Zhipu's coding plan, set "apiBase": "https://open.bigmodel.cn/api/coding/paas/v4" in your zhipu provider config.
  • Alibaba Cloud BaiLian: If you're using Alibaba Cloud BaiLian's OpenAI-compatible endpoint, set "apiBase": "https://dashscope.aliyuncs.com/compatible-mode/v1" in your dashscope provider config.
  • ModelScope: If you're using ModelScope's OpenAI-compatible endpoint, set "apiBase": "https://api-inference.modelscope.cn/v1" in your modelscope provider config.
  • StepFun Step Plan: If you're on StepFun's Step Plan subscription, set "apiBase": "https://api.stepfun.ai/step_plan/v1" in your stepfun provider config. Supported models include step-3.5-flash, step-3.5-flash-2603, and step-router-v1.
  • Step Fun (Mainland China): If your API key is from Step Fun's mainland China platform (stepfun.com), set "apiBase": "https://api.stepfun.com/v1" in your stepfun provider config.
  • Xiaomi MiMo thinking mode: MiMo models (e.g. mimo-v2.5-pro) default to enabled thinking. Use agents.defaults.reasoningEffort: "none" to disable it, or "low" / "medium" / "high" to keep it on. Omitting the field preserves the provider's per-model default.
  • Xiaomi MiMo Token Plan: If you're on MiMo's token plan, set "apiBase": "https://token-plan-sgp.xiaomimimo.com/v1" in your xiaomi_mimo provider config.
  • Custom OpenAI-compatible providers: Besides the built-in custom provider, any extra key under providers can define its own OpenAI-compatible endpoint. For example, providers.companyProxy.apiBase plus modelPresets.primary.provider: "companyProxy" creates a separate custom provider. Set apiBase. Set apiKey only when the endpoint requires it. This named-custom path uses the OpenAI-compatible request format only. For Anthropic-compatible proxies, use providers.anthropic.apiBase with provider: "anthropic".
  • Provider-scoped proxy: providers.<name>.proxy routes only that provider through an HTTP proxy. It is supported for OpenAI-compatible providers, openai_codex, and xai_grok. Native provider backends such as anthropic, bedrock, azure_openai, and github_copilot reject proxy.
ProviderPurposeGet API Key
customAny OpenAI-compatible endpoint—
openrouterLLM gateway for hosted model families + Voice transcription (STT models)openrouter.ai
edenaiLLM gateway for Eden AI's OpenAI-compatible model catalogapp.edenai.run
opencodeLLM gateway (OpenCode Zen coding-agent models)opencode.ai/docs/zen
opencode_zenLLM gateway (legacy alias for OpenCode Zen)opencode.ai/docs/zen
opencode_goLLM gateway (OpenCode Go low-cost coding models)opencode.ai/docs/go
huggingfaceLLM (Hugging Face Inference Providers)huggingface.co/settings/tokens
skyworkLLM (Skywork / APIFree API gateway)apifree.ai
volcengineLLM (VolcEngine, pay-per-use)Coding Plan · volcengine.com
volcengine_coding_planLLM (VolcEngine Coding Plan subscription endpoint)volcengine.com
byteplusLLM (VolcEngine international, pay-per-use)Coding Plan · byteplus.com
byteplus_coding_planLLM (BytePlus Coding Plan subscription endpoint)byteplus.com
anthropicLLM (Claude direct)console.anthropic.com
azure_openaiLLM (Azure OpenAI)portal.azure.com
bedrockLLM (AWS Bedrock Converse, Claude/Nova/Llama/etc.)aws.amazon.com/bedrock
openaiLLM + Voice transcription (Whisper)platform.openai.com
assemblyaiVoice transcription onlyassemblyai.com
deepseekLLM (DeepSeek direct)platform.deepseek.com
groqLLM + Voice transcription (Whisper, default)console.groq.com
minimaxLLM (MiniMax direct)platform.minimaxi.com
minimax_anthropicLLM (MiniMax Anthropic-compatible endpoint, thinking mode)platform.minimaxi.com
geminiLLM (Gemini direct)aistudio.google.com
aihubmixLLM (API gateway, access to all models)aihubmix.com
siliconflowLLM (SiliconFlow/硅基流动)siliconflow.cn
novitaLLM (Novita AI OpenAI-compatible gateway)novita.ai
dashscopeLLM (Qwen)dashscope.console.aliyun.com
modelscopeLLM (ModelScope/魔搭社区) + Image generationmodelscope.cn
moonshotLLM (Moonshot/Kimi)platform.kimi.com
kimi_codingLLM (Kimi Coding Plan, Anthropic Messages API)platform.kimi.com
zhipuLLM (Zhipu GLM)open.bigmodel.cn
xiaomi_mimoLLM (MiMo)platform.xiaomimimo.com
longcatLLM (LongCat)longcat.chat
ant_lingLLM (Ant Ling / 蚂蚁百灵)developer.ant-ling.com
ollamaLLM (local, Ollama)—
lm_studioLLM (local, LM Studio)—
atomic_chatLLM (local, Atomic Chat)—
mistralLLMdocs.mistral.ai
stepfunLLM (Step Fun/阶跃星辰) + Voice transcription (ASR)platform.stepfun.com
ovmsLLM (local, OpenVINO Model Server)docs.openvino.ai
vllmLLM (local, any OpenAI-compatible server)—
nvidiaLLM (NVIDIA NIM)build.nvidia.com
openai_codexLLM (Codex, OAuth)nanoinfra provider login openai-codex --set-main
xai_grokLLM (Grok, OAuth)nanoinfra provider login xai-grok --set-main
github_copilotLLM (GitHub Copilot, OAuth)nanoinfra provider login github-copilot
qianfanLLM (Baidu Qianfan)cloud.baidu.com
OpenAI

By default, OpenAI uses apiType: "auto": nanoinfra calls Chat Completions normally and routes GPT-5/o-series or explicit reasoningEffort requests through the Responses API when useful. You can force a specific API surface:

{
"providers": {
"openai": {
"apiKey": "${OPENAI_API_KEY}",
"apiType": "chat_completions"
}
}
}

Valid apiType values are exactly auto, chat_completions, and responses.

extraQuery adds query-string parameters to every request the provider makes. An Azure-style gateway needs that when it expects a version in the URL rather than in a header:

{
"providers": {
"custom": {
"apiBase": "https://gateway.example.com/v1",
"extraQuery": { "api-version": "2026-05-01" }
}
}
}

It is accepted on the OpenAI-compatible providers and on providers.bedrock. extraHeaders is the same idea for headers, and extraBody for request fields.

extraBody follows the selected OpenAI API surface. With Chat Completions, nanoinfra passes it through as the SDK extra_body value. With Responses, configure it in Responses API body shape. nanoinfra merges ordinary top-level fields into the Responses request body, appends extraBody.tools after generated function tools, and merges extraBody.include without duplicates:

{
"providers": {
"openai": {
"apiKey": "${OPENAI_API_KEY}",
"apiType": "responses",
"extraBody": {
"tools": [{ "type": "web_search" }],
"include": ["web_search_call.action.sources"]
}
}
}
}

Responses conversation state and compaction​

Providers that use the Responses API can keep reasoning context across a conversation, which helps with multi-step tasks. Supported providers can also compact long conversations automatically.

nanoinfra preserves Responses conversation state automatically for OpenAI Responses, OpenAI Codex, Azure OpenAI, DeepSeek V4 Flash, and compatible GitHub Copilot models. Native compaction is also automatic when the provider supports it. The threshold is derived from the active model's context window and reserved output headroom. No provider configuration is required.

Azure OpenAI

The azure_openai provider talks to your Azure OpenAI resource via the OpenAI Responses API (/openai/v1/responses). Model names map to deployment names, not OpenAI model IDs. Two authentication modes are supported.

Mode 1: Static API key (simplest)

{
"providers": {
"azure_openai": {
"apiKey": "${AZURE_OPENAI_API_KEY}",
"apiBase": "https://my-resource.openai.azure.com"
}
},
"modelPresets": {
"azure": {
"provider": "azure_openai",
"model": "my-gpt-5-deployment"
}
},
"agents": {
"defaults": {
"modelPreset": "azure"
}
}
}

Mode 2: Microsoft Entra ID (Azure AD) via DefaultAzureCredential

Omit apiKey (or leave it empty / unset). The provider falls back to DefaultAzureCredential and acquires a bearer token scoped to https://cognitiveservices.azure.com/.default for every request. The Azure SDK's own MSAL-backed cache returns valid tokens without a network round-trip.

{
"providers": {
"azure_openai": {
"apiBase": "https://my-resource.openai.azure.com"
}
},
"modelPresets": {
"azure": {
"provider": "azure_openai",
"model": "my-gpt-5-deployment"
}
},
"agents": {
"defaults": {
"modelPreset": "azure"
}
}
}

Install the optional dependency:

nanoinfra plugins enable azure

DefaultAzureCredential walks this chain in order and uses the first identity that succeeds:

  1. EnvironmentCredential — reads AZURE_TENANT_ID, AZURE_CLIENT_ID, and one of AZURE_CLIENT_SECRET / AZURE_CLIENT_CERTIFICATE_PATH / AZURE_USERNAME + AZURE_PASSWORD.
  2. WorkloadIdentityCredential — for AKS workload-identity / federated tokens (AZURE_FEDERATED_TOKEN_FILE).
  3. ManagedIdentityCredential — for Azure VMs, App Service, Functions, Container Apps, etc.
  4. AzureCliCredential — uses the token from az login on your dev machine.
  5. AzurePowerShellCredential — uses the token from Connect-AzAccount.
  6. AzureDeveloperCliCredential — uses the token from azd auth login.
  7. InteractiveBrowserCredential (disabled by default).

The identity that ends up signing the request must be assigned the Cognitive Services OpenAI User RBAC role (or higher) on the Azure OpenAI resource. Without that role you will see 401/403 errors at the first request.

apiBase remains mandatory in both modes — it's your Azure resource endpoint and cannot be inferred. If neither apiKey is set nor azure-identity is installed, the provider raises a clear error pointing you at nanoinfra plugins enable azure.

Skywork / APIFree

Skywork uses APIFree's OpenAI-compatible Agent API endpoint. Configure the provider once, then use Skywork model IDs such as skywork-ai/skyclaw-v1.

{
"providers": {
"skywork": {
"apiKey": "${SKYWORK_API_KEY}",
"apiBase": "https://api.apifree.ai/agent/v1"
}
},
"modelPresets": {
"skywork": {
"provider": "skywork",
"model": "skywork-ai/skyclaw-v1",
"maxTokens": 32768,
"contextWindowTokens": 131072
}
},
"agents": {
"defaults": {
"modelPreset": "skywork"
}
}
}

You can also reference ${APIFREE_API_KEY} in apiKey if that is how your environment names the credential.

AWS Bedrock (Converse API)

Bedrock uses the native bedrock-runtime Converse API. So it can call any Bedrock model ID that supports Converse. Examples are Claude Opus 4.7, Claude Sonnet, Amazon Nova, Meta Llama, Mistral, and Qwen. It supports normal chat, streaming, tool calling, tool results, token usage, and Bedrock error metadata.

This provider is for Bedrock's native Converse API, not Bedrock's OpenAI-compatible /openai/v1 endpoint. For OpenAI-compatible Bedrock models, you can still use custom if you specifically want that API surface.

Install Bedrock support first:

nanoinfra plugins enable bedrock

[!NOTE] If you configured Bedrock before boto3 became an optional dependency, run nanoinfra plugins enable bedrock after upgrading. Otherwise the provider will fail when it first tries to create a Bedrock client.

1. Configure credentials

Use the normal AWS credential chain (AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY, an AWS profile, or an IAM role). The IAM identity needs:

{
"Effect": "Allow",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream"
],
"Resource": "*"
}

You can also set providers.bedrock.apiKey to a Bedrock API key. nanoinfra exports it as AWS_BEARER_TOKEN_BEDROCK for the AWS SDK.

Credential options:

  • AWS CLI/default profile: leave apiKey and profile empty, then run aws configure or provide AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY.
  • Named AWS profile: set profile to a profile from ~/.aws/config or ~/.aws/credentials.
  • IAM role: on EC2/ECS/Lambda, leave apiKey and profile empty and attach a role with Bedrock permissions.
  • Bedrock API key: set apiKey or AWS_BEARER_TOKEN_BEDROCK profile can stay null.

2. Minimal config

For a non-Anthropic model such as Amazon Nova:

{
"providers": {
"bedrock": {
"region": "us-east-1"
}
},
"modelPresets": {
"bedrockNova": {
"provider": "bedrock",
"model": "bedrock/amazon.nova-lite-v1:0",
"reasoningEffort": null
}
},
"agents": {
"defaults": {
"modelPreset": "bedrockNova"
}
}
}

With a Bedrock API key:

{
"providers": {
"bedrock": {
"region": "us-east-1",
"apiKey": "${AWS_BEARER_TOKEN_BEDROCK}"
}
},
"modelPresets": {
"bedrockNova": {
"provider": "bedrock",
"model": "bedrock/amazon.nova-lite-v1:0",
"reasoningEffort": null
}
},
"agents": {
"defaults": {
"modelPreset": "bedrockNova"
}
}
}

With a named AWS profile:

{
"providers": {
"bedrock": {
"region": "us-east-1",
"profile": "my-bedrock-profile"
}
},
"modelPresets": {
"bedrockNova": {
"provider": "bedrock",
"model": "bedrock/amazon.nova-lite-v1:0"
}
},
"agents": {
"defaults": {
"modelPreset": "bedrockNova"
}
}
}

3. Claude Opus 4.7 example

{
"providers": {
"bedrock": {
"region": "us-east-1"
}
},
"modelPresets": {
"bedrockClaude": {
"provider": "bedrock",
"model": "bedrock/global.anthropic.claude-opus-4-7",
"reasoningEffort": "medium",
"maxTokens": 8192
}
},
"agents": {
"defaults": {
"modelPreset": "bedrockClaude"
}
}
}

For regional routing, use one of Bedrock's inference IDs, for example bedrock/us.anthropic.claude-opus-4-7, bedrock/eu.anthropic.claude-opus-4-7, or bedrock/jp.anthropic.claude-opus-4-7.

Claude Opus 4.7 does not accept temperature, top_p, or top_k. nanoinfra omits temperature automatically for this model. If reasoningEffort is set to low, medium, high, max, or adaptive, nanoinfra sends Bedrock's adaptive thinking parameter.

Anthropic models on Bedrock can also require Anthropic use-case registration and are subject to Anthropic-supported country/region restrictions. If Claude fails with a ValidationException about unsupported countries or regions, try a non-Anthropic Bedrock model such as Amazon Nova to verify the provider setup.

4. Model IDs

Use Bedrock model IDs or inference profile IDs with a bedrock/ prefix in nanoinfra config. nanoinfra removes the prefix before calling AWS.

Examples:

  • bedrock/amazon.nova-micro-v1:0
  • bedrock/amazon.nova-lite-v1:0
  • bedrock/global.anthropic.claude-opus-4-7
  • bedrock/us.anthropic.claude-opus-4-7
  • bedrock/openai.gpt-oss-20b-1:0
  • bedrock/meta.llama...
  • bedrock/mistral...

Check the Bedrock console for the exact model ID and region availability. Some models require cross-region inference profile IDs such as us.*, eu.*, or global.*.

5. Advanced model fields

Model-specific fields can be supplied with extraBody. nanoinfra merges it into Converse additionalModelRequestFields:

{
"providers": {
"bedrock": {
"region": "us-east-1",
"extraBody": {
"thinking": {
"type": "adaptive",
"effort": "medium",
"display": "summarized"
}
}
}
}
}

Use apiBase only for a custom Bedrock Runtime endpoint URL, such as a VPC endpoint or proxy. It is not needed for normal AWS regions.

Current scope: nanoinfra passes messages, system, inferenceConfig, toolConfig, and additionalModelRequestFields. Bedrock Prompt Management, Guardrails, serviceTier, and other top-level Converse options are not first-class config fields yet.

6. Quick checks

# For AWS credential-chain usage:
aws sts get-caller-identity

# For API-key usage:
export AWS_BEARER_TOKEN_BEDROCK="your-bedrock-api-key"
export AWS_REGION="us-east-1"

Then run:

nanoinfra agent -m "Reply with one short sentence."
OpenAI Codex (OAuth)

Codex uses OAuth instead of API keys and requires a ChatGPT Plus or Pro account. Authenticate it and make the current flagship model the active agent model with one command:

nanoinfra provider login openai-codex --set-main

Then run:

nanoinfra agent -m "Hello!"

To opt in to Codex Fast mode, merge this provider setting into config.json:

{
"providers": {
"openaiCodex": {
"extraBody": {
"service_tier": "priority"
}
}
}
}

priority is the Responses API request value used by Codex Fast mode. The setting only works for models and accounts that support Fast mode. Remove service_tier to return to standard processing. Fast mode consumes Codex credits at a higher rate. See the OpenAI Codex rate card for current details.

For proxy, remote/headless login, model-name, or config-key errors, see troubleshooting.md.

xAI Grok (OAuth)

Use an eligible X Premium / Grok subscription without putting an API key in config.json:

nanoinfra provider login xai-grok --set-main
nanoinfra agent -m "Hello from Grok."

The default model is xai-grok/grok-4.5 with a 500,000-token context window. The provider reads xAI's model catalog and includes the server-hosted x_search tool only when the selected model advertises supportsBackendSearch. Models without that capability continue normally without hosted X Search. When enabled, searches run inside xAI's Responses API and citations arrive as inline links.

This is xAI subscription OAuth, not X Developer OAuth. nanoinfra follows the public OAuth client and proxy contract used by Grok Build. The browser flow uses a random loopback callback and PKCE. The resulting token is stored in the active instance's auth/xai.json (normally ~/.nanoinfra/auth/xai.json), separately from Grok Build so rotating refresh tokens cannot invalidate one another.

To use a provider-specific proxy, merge this into config.json before login:

{
"providers": {
"xaiGrok": {
"proxy": "http://127.0.0.1:7890"
}
}
}

The proxy applies to OAuth discovery, token exchange/refresh, model-catalog lookups, and subscription model requests. Because this integration depends on xAI's public Grok Build client contract, an upstream contract change may require a nanoinfra update.

GitHub Copilot (OAuth)

GitHub Copilot uses OAuth instead of API keys. Requires a GitHub account with a plan configured. No providers.github_copilot block is needed in config.json nanoinfra provider login stores the OAuth session outside config.

For GitHub Enterprise / Copilot for Business, set the endpoint overrides you need before login:

export NANOINFRA_GITHUB_COPILOT_CLIENT_ID="your-enterprise-client-id"
export NANOINFRA_GITHUB_DEVICE_CODE_URL="https://ghe.example/login/device/code"
export NANOINFRA_GITHUB_ACCESS_TOKEN_URL="https://ghe.example/login/oauth/access_token"
export NANOINFRA_GITHUB_USER_URL="https://api.ghe.example/user"
export NANOINFRA_COPILOT_TOKEN_URL="https://api.ghe.example/copilot_internal/v2/token"
export NANOINFRA_COPILOT_BASE_URL="https://copilot-api.ghe.example"

1. Login:

nanoinfra provider login github-copilot

2. Set model (merge into ~/.nanoinfra/config.json):

{
"modelPresets": {
"copilot": {
"provider": "github_copilot",
"model": "github-copilot/gpt-4.1"
}
},
"agents": {
"defaults": {
"modelPreset": "copilot"
}
}
}

3. Chat:

nanoinfra agent -m "Hello!"

# Target a specific workspace/config locally
nanoinfra agent -c ~/.nanoinfra-telegram/config.json -m "Hello!"

# One-off workspace override on top of that config
nanoinfra agent -c ~/.nanoinfra-telegram/config.json -w /tmp/nanoinfra-telegram-test -m "Hello!"

Docker users: use docker run -it for interactive OAuth login.

OpenCode Zen / Go

OpenCode Zen and OpenCode Go are available through nanoinfra's built-in OpenAI-compatible provider flow. They share the OPENCODE_API_KEY environment variable, but use separate provider keys and default base URLs:

ProviderDefault API baseModel prefix accepted by nanoinfra
opencodehttps://opencode.ai/zen/v1opencode/<model-id>
opencode_zenhttps://opencode.ai/zen/v1opencode/<model-id>
opencode_gohttps://opencode.ai/zen/go/v1opencode-go/<model-id>

OpenCode Zen:

{
"providers": {
"opencode": {
"apiKey": "${OPENCODE_API_KEY}"
}
},
"modelPresets": {
"opencodeZen": {
"provider": "opencode",
"model": "opencode/deepseek-v4-pro"
}
},
"agents": {
"defaults": {
"modelPreset": "opencodeZen"
}
}
}

providers.opencodeZen / provider: "opencode_zen" still work as compatibility aliases for existing configs.

OpenCode Go:

{
"providers": {
"opencodeGo": {
"apiKey": "${OPENCODE_API_KEY}"
}
},
"modelPresets": {
"opencodeGo": {
"provider": "opencode_go",
"model": "opencode-go/deepseek-v4-flash"
}
},
"agents": {
"defaults": {
"modelPreset": "opencodeGo"
}
}
}

OpenCode's own docs list models across responses, messages, provider-specific model endpoints, and chat/completions. nanoinfra's OpenCode providers use the OpenAI-compatible chat/completions path, so pick model IDs from that endpoint family. The opencode/... and opencode-go/... prefixes are accepted for config readability and stripped before sending the request.

LongCat (OpenAI-compatible)

LongCat is available through nanoinfra's built-in OpenAI-compatible provider flow. The default API base already points to https://api.longcat.chat/openai/v1, so you usually only need to set apiKey.

{
"providers": {
"longcat": {
"apiKey": "${LONGCAT_API_KEY}"
}
},
"modelPresets": {
"longcat": {
"provider": "longcat",
"model": "LongCat-2.0-Preview",
"maxTokens": 8192,
"contextWindowTokens": 1048576
}
},
"agents": {
"defaults": {
"modelPreset": "longcat"
}
}
}

Current LongCat API docs list LongCat-2.0-Preview as the supported model. The older LongCat-Flash-* models were retired by LongCat on 2026-05-29.

Xiaomi MiMo

Xiaomi MiMo models are automatically detected by the xiaomi_mimo provider when the model name contains mimo. The default API base is https://api.xiaomimimo.com/v1.

Token Plan: If you're using MiMo's token plan, override apiBase with the dedicated endpoint:

{
"providers": {
"xiaomi_mimo": {
"apiKey": "${XIAOMIMIMO_API_KEY}",
"apiBase": "https://token-plan-sgp.xiaomimimo.com/v1"
}
},
"modelPresets": {
"mimo": {
"provider": "xiaomi_mimo",
"model": "xiaomi/mimo-v2.5-pro"
}
},
"agents": {
"defaults": {
"modelPreset": "mimo"
}
}
}

Use the model ID and API key from the MiMo token plan console, and check the MiMo platform for the latest supported model names.

StepFun Step Plan (subscription)

Step Plan is StepFun's subscription-based service for high-frequency AI developers. If you're on a Step Plan subscription, override apiBase in the existing stepfun provider config to point to the dedicated Step Plan endpoint.

{
"providers": {
"stepfun": {
"apiKey": "${STEPFUN_API_KEY}",
"apiBase": "https://api.stepfun.ai/step_plan/v1"
}
},
"modelPresets": {
"stepfun": {
"provider": "stepfun",
"model": "step-3.5-flash"
}
},
"agents": {
"defaults": {
"modelPreset": "stepfun"
}
}
}

Supported models include step-3.5-flash, step-3.5-flash-2603, and step-router-v1.

Ant Ling (OpenAI-compatible)

Ant Ling is available through nanoinfra's built-in OpenAI-compatible provider flow. The default API base points to https://api.ant-ling.com/v1, so you usually only need to set apiKey.

{
"providers": {
"antLing": {
"apiKey": "${ANT_LING_API_KEY}"
}
},
"modelPresets": {
"antLing": {
"provider": "ant_ling",
"model": "Ling-2.6-flash"
}
},
"agents": {
"defaults": {
"modelPreset": "antLing"
}
}
}

Official OpenAI-compatible model names include Ling-2.6-1T, Ling-2.6-flash, Ling-2.5-1T, Ling-1T, Ring-2.5-1T, and Ring-1T.

Custom Provider (Any OpenAI-compatible API)

Connects directly to any OpenAI-compatible endpoint — llama.cpp, Together AI, Fireworks, Azure OpenAI, or any self-hosted server. Model name is passed as-is.

{
"providers": {
"custom": {
"apiKey": "your-api-key",
"apiBase": "https://api.your-provider.com/v1"
}
},
"modelPresets": {
"custom": {
"provider": "custom",
"model": "your-model-name"
}
},
"agents": {
"defaults": {
"modelPreset": "custom"
}
}
}

For local servers that don't require authentication, set apiKey to null.

custom is the right choice for providers that expose an OpenAI-compatible chat completions API. It does not force third-party endpoints onto the OpenAI/Azure Responses API.

If your proxy or gateway is specifically Responses-API-compatible, configure the azure_openai provider shape and point apiBase at that endpoint:

{
"providers": {
"azure_openai": {
"apiKey": "your-api-key",
"apiBase": "https://api.your-provider.com",
"defaultModel": "your-model-name"
}
},
"modelPresets": {
"responsesProxy": {
"provider": "azure_openai",
"model": "your-model-name"
}
},
"agents": {
"defaults": {
"modelPreset": "responsesProxy"
}
}
}

Anthropic-compatible endpoints are separate: use providers.anthropic.apiBase and set the preset provider to anthropic. Arbitrary custom provider names do not use the Anthropic Messages API format.

In short:

  • chat-completions-compatible endpoint → custom or a named custom provider.
  • Responses-compatible endpoint → azure_openai.
  • Anthropic-compatible endpoint → anthropic with apiBase.

Some OpenAI-compatible gateways expose request-body extensions such as vLLM guided decoding or local sampling controls. Put those under extraBody. nanoinfra merges them into the chat-completions request body after its provider defaults:

{
"providers": {
"custom": {
"apiKey": "your-api-key",
"apiBase": "https://api.your-provider.com/v1",
"extraBody": {
"repetition_penalty": 1.15,
"chat_template_kwargs": {
"enable_thinking": false
}
}
}
}
}

If a custom OpenAI-compatible endpoint exposes a provider-specific thinking toggle, set thinkingStyle so nanoinfra can translate reasoningEffort into the right request body. Supported styles are thinking_type ({"thinking":{"type":"enabled"}}), enable_thinking ({"enable_thinking": true}), and reasoning_split ({"reasoning_split": true}):

{
"providers": {
"companyProxy": {
"apiKey": "${COMPANY_PROXY_API_KEY}",
"apiBase": "https://api.your-provider.com/v1",
"thinkingStyle": "enable_thinking"
}
},
"modelPresets": {
"company": {
"provider": "companyProxy",
"model": "served-model-name",
"reasoningEffort": "high"
}
}
}

Leave thinkingStyle unset unless the endpoint explicitly documents one of those wire formats. extraBody is still applied last, so advanced users can override the generated value.

Ollama (local)

Run a local model with Ollama, then add to config:

1. Start Ollama (example):

ollama run llama3.2

2. Add to config (partial — merge into ~/.nanoinfra/config.json):

{
"providers": {
"ollama": {
"apiBase": "http://localhost:11434/v1"
}
},
"modelPresets": {
"ollama": {
"provider": "ollama",
"model": "llama3.2"
}
},
"agents": {
"defaults": {
"modelPreset": "ollama"
}
}
}

provider: "auto" also works when providers.ollama.apiBase is configured, but pinning "provider": "ollama" inside the preset is the clearest option.

LM Studio (local)

LM Studio provides a local OpenAI-compatible server for running LLMs. Download models through the LM Studio UI, then start the local server.

1. Start LM Studio server:

  • Launch LM Studio.
  • Go to the "Local Server" tab.
  • Load a model (e.g., Llama, Mistral, Qwen).
  • Click "Start Server" (default port: 1234).

2. Add to config (partial — merge into ~/.nanoinfra/config.json):

{
"providers": {
"lm_studio": {
"apiKey": null,
"apiBase": "http://localhost:1234/v1"
}
},
"modelPresets": {
"lmStudio": {
"provider": "lm_studio",
"model": "local-model"
}
},
"agents": {
"defaults": {
"modelPreset": "lmStudio"
}
}
}

Note: Set apiKey to null for LM Studio since it runs locally and doesn't require authentication. The model name should match what's shown in the LM Studio UI. provider: "auto" also works when providers.lm_studio.apiBase is configured, but pinning "provider": "lm_studio" inside the preset is the clearest option.

Atomic Chat (local)

Atomic Chat is a local-first desktop app that exposes an OpenAI-compatible HTTP API (default http://localhost:1337/v1). This setup applies when you want to run nanoinfra against a model on your own machine instead of a hosted API provider.

1. Start Atomic Chat

  • Install Atomic Chat on your machine.
  • Open Atomic Chat, download a model, and keep the app running. The local API is enabled by default.
  • Copy the model ID exposed by the local API. For example, the model ID for Qwen 3 32B might be qwen3-32b.

2. Add to config (partial — merge into ~/.nanoinfra/config.json):

{
"providers": {
"atomic_chat": {
"apiKey": null,
"apiBase": "http://localhost:1337/v1"
}
},
"modelPresets": {
"atomic": {
"provider": "atomic_chat",
"model": "qwen3-32b"
}
},
"agents": {
"defaults": {
"modelPreset": "atomic"
}
}
}

Note: Replace qwen3-32b with the model ID from Atomic Chat. Set apiKey to null if your Atomic Chat server does not require a key. If it does, set apiKey (or the ATOMIC_CHAT_API_KEY environment variable) to the value Atomic Chat expects.

provider: "auto" also works when providers.atomic_chat.apiBase is configured, but pinning "provider": "atomic_chat" inside the preset is the clearest option.

OpenVINO Model Server (local / OpenAI-compatible)

Run LLMs locally on Intel GPUs using OpenVINO Model Server. OVMS exposes an OpenAI-compatible API at /v3.

Requires Docker and an Intel GPU with driver access (/dev/dri).

1. Pull the model (example):

mkdir -p ov/models && cd ov

docker run -d \
--rm \
--user $(id -u):$(id -g) \
-v $(pwd)/models:/models \
openvino/model_server:latest-gpu \
--pull \
--model_name openai/gpt-oss-20b \
--model_repository_path /models \
--source_model OpenVINO/gpt-oss-20b-int4-ov \
--task text_generation \
--tool_parser gptoss \
--reasoning_parser gptoss \
--enable_prefix_caching true \
--target_device GPU

This downloads the model weights. Wait for the container to finish before proceeding.

2. Start the server (example):

docker run -d \
--rm \
--name ovms \
--user $(id -u):$(id -g) \
-p 8000:8000 \
-v $(pwd)/models:/models \
--device /dev/dri \
--group-add=$(stat -c "%g" /dev/dri/render* | head -n 1) \
openvino/model_server:latest-gpu \
--rest_port 8000 \
--model_name openai/gpt-oss-20b \
--model_repository_path /models \
--source_model OpenVINO/gpt-oss-20b-int4-ov \
--task text_generation \
--tool_parser gptoss \
--reasoning_parser gptoss \
--enable_prefix_caching true \
--target_device GPU

3. Add to config (partial — merge into ~/.nanoinfra/config.json):

{
"providers": {
"ovms": {
"apiBase": "http://localhost:8000/v3"
}
},
"modelPresets": {
"ovms": {
"provider": "ovms",
"model": "openai/gpt-oss-20b"
}
},
"agents": {
"defaults": {
"modelPreset": "ovms"
}
}
}

OVMS is a local server — no API key required. Supports tool calling (--tool_parser gptoss), reasoning (--reasoning_parser gptoss), and streaming. See the official OVMS docs for more details.

vLLM (local / OpenAI-compatible)

Run your own model with vLLM or any OpenAI-compatible server, then add to config:

1. Start the server (example):

vllm serve meta-llama/Llama-3.1-8B-Instruct --port 8000

2. Add to config (partial — merge into ~/.nanoinfra/config.json):

Provider (set API key to null for local servers):

{
"providers": {
"vllm": {
"apiKey": null,
"apiBase": "http://localhost:8000/v1"
}
}
}

Model preset:

{
"modelPresets": {
"vllm": {
"provider": "vllm",
"model": "meta-llama/Llama-3.1-8B-Instruct"
}
},
"agents": {
"defaults": {
"modelPreset": "vllm"
}
}
}

Contributor notes for adding new providers live in development.md.

Model Presets​

Model presets let you name a complete model configuration and select one per session with /model <preset>. They are the recommended way to configure models because the same names can be reused for new-session defaults, chat-command switching, and fallback chains.

Existing configs do not need to change. Direct agents.defaults.model, provider, maxTokens, contextWindowTokens, temperature, and reasoningEffort fields still define the implicit default preset. For new configs, prefer top-level modelPresets plus agents.defaults.modelPreset.

Nothing ships a model. agents.defaults.model is empty until you set one. A fresh install has no model, and says so rather than appearing to run on one it has no credential for. The first modelPresets entry you add is the primary one: it answers until you choose another, and adding it through the WebUI activates it. A deployment with no model and no presets refuses every turn with a line telling you to add one.

{
"modelPresets": {
"fast": {
"label": "Fast",
"model": "gpt-4.1-mini",
"provider": "openai",
"maxTokens": 4096,
"contextWindowTokens": 128000,
"temperature": 0.2,
"reasoningEffort": "low"
},
"deep": {
"label": "Deep",
"model": "claude-opus-4-5",
"provider": "anthropic",
"maxTokens": 8192,
"contextWindowTokens": 200000,
"reasoningEffort": "high"
},
"localSmall": {
"label": "Local Small",
"model": "llama3.2",
"provider": "ollama",
"maxTokens": 4096,
"contextWindowTokens": 32768,
"temperature": 0.2
}
},
"agents": {
"defaults": {
"modelPreset": "fast",
"fallbackModels": ["deep", "localSmall"]
}
}
}

modelPresets is a top-level object. The keys under it (fast, deep, coding, etc.) are user-defined preset names. Each preset supports:

FieldDescription
labelOptional display name shown in model lists.
modelModel name to use for this preset.
providerProvider name, or "auto" to use provider auto-detection.
maxTokensMaximum completion/output tokens.
contextWindowTokensContext window size used by prompt building and consolidation decisions.
temperatureSampling temperature.
reasoningEffortOptional reasoning/thinking setting. Provider support varies.

default is reserved and always means the implicit preset built from direct agents.defaults.* fields. Do not define modelPresets.default. Use /model default to switch back to those direct fields in an existing config.

Set agents.defaults.modelPreset to choose the preset followed by sessions that have no saved model selection. When modelPreset is null or omitted, such sessions follow the implicit default preset from direct agents.defaults.* fields. /model <preset> saves an override in the current session, so its future turns keep that preset across process restarts while other sessions remain unchanged. The command does not write the selection back to config.json.

Model Fallbacks​

agents.defaults.fallbackModels defines an ordered failover chain for the active model configuration. The primary model is still selected by agents.defaults.modelPreset or, in older configs, by the implicit default preset from direct agents.defaults.* fields.

Each fallback candidate can be either:

  • A preset name from modelPresets, such as "deep". This is the recommended form. The preset's full model, provider, generation, and context-window config is used.
  • An inline fallback object with at least provider and model. Optional maxTokens, contextWindowTokens, and temperature fields inherit from the active primary config when omitted. reasoningEffort does not inherit. Omit it to leave reasoning off for that fallback, or set it explicitly for models that support reasoning.

Preset fallback chain:

{
"modelPresets": {
"fast": {
"model": "gpt-4.1-mini",
"provider": "openai",
"maxTokens": 4096,
"contextWindowTokens": 128000,
"temperature": 0.2
},
"deep": {
"model": "claude-opus-4-5",
"provider": "anthropic",
"maxTokens": 8192,
"contextWindowTokens": 200000,
"reasoningEffort": "high"
},
"localSmall": {
"model": "llama3.2",
"provider": "ollama",
"maxTokens": 4096,
"contextWindowTokens": 32768
}
},
"agents": {
"defaults": {
"modelPreset": "fast",
"fallbackModels": ["deep", "localSmall"]
}
}
}

String entries are preset names, not raw model names. In the example above, "deep" means modelPresets.deep. nanoinfra will not interpret it as a provider model ID. Changing a preset updates both /model <preset> switching and any fallback chain that references it.

Inline fallback object:

{
"modelPresets": {
"fast": {
"provider": "openrouter",
"model": "anthropic/claude-sonnet-4.5",
"maxTokens": 4096,
"contextWindowTokens": 65536
}
},
"agents": {
"defaults": {
"modelPreset": "fast",
"fallbackModels": [
{
"provider": "deepseek",
"model": "deepseek-v4-pro",
"maxTokens": 4096,
"contextWindowTokens": 262144
}
]
}
}
}

Use inline objects only when a fallback is not worth naming as a reusable preset. fallbackModels belongs under agents.defaults, not inside individual modelPresets entries.

Failover normally runs when the primary provider returns a fallbackable model/provider error before any answer text has been streamed. Stream-stall timeouts are the recovery exception. If the provider already emitted partial answer text and then stalls, nanoinfra closes the current stream segment. It then retries, or fails over, in a new segment. Typical fallback cases include timeouts, connection errors, 5xx server errors, 429 rate limits, overloads, authentication/permission failures such as invalid or expired credentials, and quota/balance exhaustion. It does not run for malformed requests, content filtering/refusals, or context-length/message-format errors.

If fallback candidates use smaller contextWindowTokens values, nanoinfra builds context using the smallest window in the active chain so every candidate can receive the same prompt.

Transcription Settings​

Audio transcription is a shared capability used by chat-channel voice messages and by WebUI microphone input. Chat-channel voice messages are transcribed automatically before they enter the agent. WebUI microphone input is transcribed into the composer first, so you can edit the text before sending.

Configure transcription under the top-level transcription section:

{
"transcription": {
"enabled": true,
"provider": "groq",
"model": null,
"language": null,
"maxDurationSec": 120,
"maxUploadMb": 25
}
}
SettingDefaultDescription
enabledtrueEnables audio transcription for both chat-channel voice messages and WebUI microphone input.
provider"groq"Transcription backend: "groq", "openai", "openrouter", "xiaomi_mimo", "stepfun", or "assemblyai".
modelprovider defaultOptional transcription model override. Defaults to whisper-large-v3 for Groq, whisper-1 for OpenAI, openai/whisper-1 for OpenRouter, mimo-v2.5-asr for Xiaomi MiMo ASR, stepaudio-2.5-asr for StepFun ASR, and universal-3-pro,universal-2 for AssemblyAI. OpenRouter accepts only speech-to-text models on its transcription endpoint, such as nvidia/parakeet-tdt-0.6b-v3, openai/whisper-1, or openai/gpt-4o-transcribe. Chat LLMs are rejected there. AssemblyAI accepts a comma-separated model fallback list.
languagenullOptional ISO-639 language hint, e.g. "en", "zh", "ko", or "ja".
maxDurationSec120Maximum WebUI recording duration.
maxUploadMb25Maximum WebUI audio upload size.

Provider and language resolution is intentionally ordered for backwards compatibility:

  1. transcription.provider / transcription.language
  2. Legacy channels.transcriptionProvider / channels.transcriptionLanguage
  3. Built-in defaults (provider: "groq", no language hint)

The legacy channels.* transcription fields existed before transcription became a shared capability across chat channels and WebUI microphone input. They are still read so older config.json files keep working, but they are no longer the preferred configuration surface. If both old and new fields are present, the top-level transcription values are the source of truth.

Transcription credentials are intentionally not stored in transcription. Put the API key and optional endpoint in the matching provider config:

{
"providers": {
"groq": {
"apiKey": "gsk-...",
"apiBase": "https://api.groq.com/openai/v1"
}
},
"transcription": {
"provider": "groq",
"language": "zh"
}
}

Selecting a transcription provider does not configure credentials by itself. For example, the effective provider may default to Groq for compatibility, but transcription is only usable when providers.groq.apiKey or the matching environment-backed config is available. The Settings UI writes only the top-level transcription fields.

If you are adding a new transcription provider, see development.md.