Skip to main content

Configuration

This page is one of four. The reference used to be a single 2,929-line document whose own second section was a hand-written table of contents. It was a page admitting it could not be navigated. It is split by subject now, so the fields for a thing sit beside the pages that explain the thing:

PageHolds
Configurationhow config is loaded, secrets, environment variables, channels, and the deployment-wide settings
Provider and Model Configurationevery provider, model preset, fallback and transcription field
Agent and Tool Configurationnamed agents, tool groups, web tools, MCP, knowledge and connectors
Security Configurationcapability gates, approvers, standing grants and pairing

Config file: ~/.nanoinfra/config.json

This is the full reference. If this is your first install, start with quick-start.md. If you are trying to choose a model or fix provider/model matching, use providers.md first. Come back here for exact fields and advanced options.

For normal local use, prefer the WebUI before editing JSON:

  • Settings → Models manages model choices and provider credentials.
  • Settings → Channels guides chat-platform setup.
  • Other Settings pages cover built-in capabilities.
  • Apps manages CLI App and MCP integrations.

Edit config.json directly when you need an advanced field, automate deployment, or intentionally manage configuration as code.

The JSON examples below are usually partial snippets to merge into your existing config, not full replacement files. For the mental model behind config, workspace, gateway, channels, sessions, tools, and memory, see concepts.md.

The generated config.json uses camelCase keys such as apiKey and intervalS. snake_case keys are also accepted for compatibility, but the docs prefer camelCase because that is what nanoinfra writes back to disk.

For setup and runtime failures, follow the diagnosis order in troubleshooting.md before changing multiple config areas at once.

[!NOTE] If your config file is older than the current schema, run nanoinfra onboard --refresh. nanoinfra adds missing default fields while preserving your existing values.

Where a Setting Lives​

If the WebUI does not expose the option you need, start from the task below. Most advanced changes touch one config section and one verification command.

TaskFirst keys to checkVerify withDeep dive
Make the first model reply workproviders.<name>.apiKey, optional providers.<name>.apiBase, modelPresets.<preset>, agents.defaults.modelPresetnanoinfra status, then nanoinfra agent -m "Hello!"Providers, Model Presets
Add fallback modelsmodelPresets.<fallback>, agents.defaults.fallbackModelsnanoinfra status, then a normal agent runModel Fallbacks
Keep secrets out of the config file${ENV_VAR} placeholders inside any string valueStart nanoinfra from the same environment that sets the variableEnvironment Variables for Secrets
Open the bundled WebUIchannels.websocket.enabled, optional channels.websocket.port, channels.websocket.tokenIssueSecretnanoinfra webuiChannel Settings, WebSocket docs
Connect one chat appchannels.<channel>.enabled, channel credentials, optional pairing or channels.<channel>.allowFromnanoinfra channels status, then nanoinfra gateway --verboseChannel Settings, Channels
Enable voice transcriptiontranscription.enabled, transcription.provider, matching providers.<name>.apiKeySend or upload a short voice message through a configured surfaceTranscription Settings
Enable web search or fetchtools.web.search.*, tools.web.fetch.*, optional tools.ssrfWhitelistAsk a question that requires current web information, then inspect logs if neededWeb Tools, Security
Enable image generationtools.imageGeneration.enabled, tools.imageGeneration.provider, tools.imageGeneration.model, matching provider credentialsEnable Image Generation in the WebUI and send one image requestImage Generation
Add external tools through MCPtools.mcpServers.<name>Start nanoinfra gateway --verbose and check startup/tool logsMCP
Tighten tool and network safetytools.restrictToWorkspace, tools.exec.sandbox, tools.ssrfWhitelist, channels.*.allowFromRun the same workflow through the channel or CLI you plan to exposeSecurity, Pairing
Let an automation run a remote command againgates.unattended.mutate.remote, gates.standingGrantsRead the gates: line that nanoinfra gateway logs at startCapability Gates, Capability gate reference
Tune request timeouts or process concurrencyNANOINFRA_LLM_TIMEOUT_S, NANOINFRA_STREAM_IDLE_TIMEOUT_S, NANOINFRA_MAX_CONCURRENT_REQUESTSStart nanoinfra from the same environment and inspect startup/runtime logsRuntime Environment Variables
Run multiple isolated botsseparate --config and --workspace paths, plus distinct gateway.port or channel ports when processes run togetherUse the same explicit paths with nanoinfra status, agent, webui, gateway, and serveMultiple Instances, CLI Reference
Observe model callsLANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, LANGFUSE_BASE_URL environment variablesRun one model call, then check the matching Langfuse projectLangfuse Observability
See spend in money, not tokenspricing["<provider>/<model>"].inputPerMtok and the three other ratesOpen Metrics → Usage and check the cost card names a figure rather than "No prices configured"Metrics and Pricing
Scrape the runtime gaugesgateway.metricsEnabled, gateway.metricsTokencurl -H "Authorization: Bearer <token>" http://127.0.0.1:18790/metricsMetrics and Pricing

Environment Variables for Secrets​

Instead of storing secrets directly in config.json, you can use ${VAR_NAME} references that are resolved from environment variables at startup:

{
"channels": {
"telegram": { "token": "${TELEGRAM_TOKEN}" },
"email": {
"imapPassword": "${IMAP_PASSWORD}",
"smtpPassword": "${SMTP_PASSWORD}"
}
},
"providers": {
"groq": { "apiKey": "${GROQ_API_KEY}" }
}
}

Any string value in config.json can use ${VAR_NAME}. Resolution runs once at startup, in memory only. Resolved values are never written back to disk, so editing config through nanoinfra onboard or the WebUI preserves the placeholder.

If a referenced variable is unset, nanoinfra fails fast and reports the exact config field and variable name without echoing the field value. Run nanoinfra status with the same --config path to inspect the problem.

More examples​

MCP servers — both stdio env and HTTP headers:

{
"tools": {
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" }
},
"remote": {
"url": "https://example.com/mcp/",
"headers": { "Authorization": "Bearer ${REMOTE_MCP_TOKEN}" }
}
}
}
}

Web search providers:

{
"tools": {
"web": {
"search": {
"provider": "brave",
"apiKey": "${BRAVE_API_KEY}"
}
}
}
}

Loading variables at startup​

Pick whatever fits your deployment — nanoinfra only reads os.environ at startup, so any mechanism that populates the process environment works.

systemd — use EnvironmentFile= in the service unit to load variables from a file that only the deploying user can read:

# /etc/systemd/system/nanoinfra.service (excerpt)
[Service]
EnvironmentFile=/home/youruser/nanoinfra_secrets.env
User=nanoinfra
ExecStart=...
# /home/youruser/nanoinfra_secrets.env (mode 600, owned by youruser)
TELEGRAM_TOKEN=your-token-here
IMAP_PASSWORD=your-password-here

Docker — pass an env file to the locally built image (one KEY=VALUE per line), or use -e KEY=value:

docker run --rm --env-file=./nanoinfra.env \
-v ~/.nanoinfra:/home/nanoinfra/.nanoinfra \
nanoinfra agent -m "Hello"

direnv — drop a .envrc in your working directory and run direnv allow:

# .envrc (auto-loaded by direnv)
export TELEGRAM_TOKEN=your-token-here
export ANTHROPIC_API_KEY=...

Secret managers (1Password, Bitwarden, pass) — wrap the process so secrets only exist as env vars for the lifetime of the run, never on disk:

# 1Password — references in .env.tpl look like `op://Vault/Item/field`
op run --env-file=.env.tpl -- nanoinfra agent

# pass (passwordstore.org)
ANTHROPIC_API_KEY="$(pass show api/anthropic)" nanoinfra agent

# Bitwarden
ANTHROPIC_API_KEY="$(bw get password api/anthropic)" nanoinfra agent

Runtime Environment Variables​

These variables are process-level switches. Set them in the same terminal, service unit, container, or supervisor that starts nanoinfra.

Runtime controls​

VariableDefaultDescription
NANOINFRA_MAX_CONCURRENT_REQUESTS3Maximum concurrently running inbound agent requests. Must be an integer. Set 0 or a negative value for unlimited.
NANOINFRA_LLM_TIMEOUT_S300Wall-clock timeout, in seconds. Ordinary requests use this value. Streaming requests use the greater of 300 seconds or twice this value. Set 0 to disable. Sustained-goal turns bypass this wall-clock cap.
NANOINFRA_STREAM_IDLE_TIMEOUT_S90Streaming idle timeout, in seconds, used by streaming providers. Invalid or non-positive values are ignored. Values above 3600 are clamped.
NANOINFRA_OPENAI_COMPAT_TIMEOUT_S120HTTP request timeout, in seconds, for OpenAI-compatible providers. Invalid or non-positive values are ignored.
NANOINFRA_WORKSPACE_SANDBOX_ENFORCEDunsetMarks that an external workspace sandbox is already enforced. Truthy values (1, true, yes, on, enabled) use NANOINFRA_WORKSPACE_SANDBOX_PROVIDER as the label. Any other non-false value is treated as the provider name.
NANOINFRA_WORKSPACE_SANDBOX_PROVIDERunknownDisplay label for the external workspace sandbox when NANOINFRA_WORKSPACE_SANDBOX_ENFORCED is truthy, for example macos_app_sandbox or bwrap.
NANOINFRA_SANDBOX_ENFORCEDunsetLegacy compatibility alias for NANOINFRA_WORKSPACE_SANDBOX_ENFORCED.
NANOINFRA_TMUX_SOCKET_DIR${TMPDIR:-/tmp}/nanoinfra-tmux-socketsSocket directory used by the bundled tmux skill scripts.

Secrets module​

VariableDefaultDescription
NANOINFRA_SECRETS_KEYunsetMaster encryption key (Fernet) for the Secrets module. nanoinfra never generates or stores this key — you generate it yourself and set it in the same environment that starts the gateway (service unit, container, or .env): python3 -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())". Without it, the gateway still starts fine. Every Secrets operation just returns a clean "not configured" error until it's set.
NANOINFRA_SECRETS_POSTGRES_DSNunsetPostgres connection string for secrets stored with providerId="postgres" (shared across a deployment). Optional — providerId="local" secrets work fully without it. Only Postgres-backed secrets fail with "not configured" if it's unset. Requires the secrets-postgres extra: nanoinfra plugins enable secrets-postgres (or pip install nanoinfra[secrets-postgres]).

Servers module (execution backends)​

No dedicated environment variables — execute_on_server resolves its credential from the Secrets module (above) and connects via whichever provider a Server record specifies. Each remote-connection provider needs its own optional library, all bundled under one servers extra: nanoinfra plugins enable servers (or pip install nanoinfra[servers]) installs asyncssh (the ssh provider), ansible-runner (the ansible-runner provider), and boto3 (the ssm provider). The api provider needs nothing extra — it uses the already-core httpx. A missing library only breaks the one provider that needs it. The gateway starts fine either way, and inventory CRUD (list_servers/get_server/create_server/update_server/delete_server) never needs any of them.

Installer, build, and WebUI development​

VariableDefaultDescription
NANOINFRA_BIN_DIR$HOME/.local/binInstaller launcher directory on macOS/Linux.
NANOINFRA_VENV$HOME/.nanoinfra/venvManaged virtual environment path used by the installer fallback.
NANOINFRA_SKIP_WIZARDunsetSet to 1 to skip the automatic WebUI or wizard setup a first run would start.
NANOINFRA_SKIP_WEBUI_BUILDunsetSet to 1 to skip bundling the WebUI during package builds.
NANOINFRA_FORCE_WEBUI_BUILDunsetSet to 1 to rebuild the bundled WebUI even when nanoinfra/web/dist/index.html already exists.
NANOINFRA_EXTRASunsetDocker build argument containing comma-separated Python extras such as bedrock.
NANOINFRA_CHANNELSwhatsappDocker build argument containing comma-separated channels whose manifest dependencies are preinstalled.
NANOINFRA_API_URLhttp://127.0.0.1:8765Gateway target for the Vite WebUI dev server proxy.

Internal variables such as NANOINFRA_RESTART_* and NANOINFRA_PATH_* are set by nanoinfra itself and are not a supported user configuration surface.

Channel Settings​

Global settings that apply to all channels. Configure under the channels section in ~/.nanoinfra/config.json:

{
"channels": {
"sendProgress": true,
"sendToolHints": true,
"sendMaxRetries": 3,
"telegram": {
"enabled": false
}
}
}
SettingDefaultDescription
sendProgresstrueStream agent's text progress to the channel
sendToolHintstrueStream tool-call hints (e.g. read_file("…"))
showReasoningtrueAllow channels to surface model reasoning/thinking content (DeepSeek-R1 reasoning_content, Anthropic thinking_blocks, inline <think> tags). Reasoning flows as a dedicated stream with _reasoning_delta / _reasoning_end markers — channels override send_reasoning_delta / send_reasoning_end to render in-place updates. Even with true, channels without those overrides stay no-op silently. Currently surfaced on CLI and WebSocket/WebUI (italic shimmer header, auto-collapses after the stream ends). Telegram / Slack / Discord / Matrix / Mattermost keep the base no-op until their bubble UI is adapted. Independent of sendProgress.
sendMaxRetries3Max delivery attempts per outbound message, including the initial send (0-10 configured, minimum 1 actual attempt)

Non-image attachments are included in the user message as local path references, without injecting their contents into the model prompt. When file tools are enabled, the agent can inspect supported text, PDF, DOCX, XLSX, and PPTX files on demand with read_file. It can also pass the original path to another tool when exact file bytes are required. The deprecated channels.extractDocumentText setting is accepted for compatibility but ignored. Normal tool workspace and media access rules still apply to attachment paths.

channels.transcriptionProvider and channels.transcriptionLanguage are deprecated compatibility fields. They remain as a read-only fallback for older configs, but new configuration should use top-level transcription.provider and transcription.language.

sendProgress and sendToolHints can also be overridden per channel. The global values stay as defaults for channels that do not set their own value:

{
"channels": {
"sendProgress": true,
"sendToolHints": true,
"telegram": {
"enabled": true,
"sendProgress": false,
"sendToolHints": false
},
"websocket": {
"enabled": true,
"sendToolHints": true
}
}
}

Telegram richMessages defaults to false. Enable it only to opt in to Bot API 10.1 sendRichMessage rendering. Leave it disabled for Telegram Web clients that show unsupported-message errors for rich messages.

Retry Behavior​

Retry is intentionally simple.

When a channel send() raises, nanoinfra retries at the channel-manager layer. By default, channels.sendMaxRetries is 3, and that count includes the initial send.

  • Attempt 1: Send immediately.
  • Attempt 2: Retry after 1s.
  • Attempt 3: Retry after 2s.
  • Higher retry budgets: Backoff continues as 1s, 2s, 4s, then stays capped at 4s.
  • Transient failures: Network hiccups and temporary API limits often recover on the next attempt.
  • Permanent failures: Invalid tokens, revoked access, or banned channels will exhaust the retry budget and fail cleanly.

[!NOTE] This design is deliberate: channel implementations should raise on delivery failure, and the channel manager owns the shared retry policy.

Some channels may still apply small API-specific retries internally. For example, Telegram separately retries timeout and flood-control errors before surfacing a final failure to the manager.

If a channel is completely unreachable, nanoinfra cannot notify the user through that same channel. Watch logs for Failed to send to {channel} after N attempts to spot persistent delivery failures.

Langfuse Observability​

nanoinfra can trace OpenAI-compatible provider calls through Langfuse's OpenAI SDK wrapper. This is configured with environment variables, not config.json.

Install the optional package in the same Python environment that runs nanoinfra:

nanoinfra plugins enable langfuse

Set Langfuse credentials before starting nanoinfra agent, nanoinfra gateway, or nanoinfra serve:

export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"

When LANGFUSE_SECRET_KEY is set and the langfuse package is installed, nanoinfra uses langfuse.openai.AsyncOpenAI for OpenAI-compatible providers so model requests are sent to Langfuse in the background. If the secret key is set but langfuse is missing, nanoinfra logs a warning and falls back to the regular OpenAI client.

Use the Langfuse region or self-hosted URL that matches your project. The Langfuse OpenAI SDK docs use LANGFUSE_BASE_URL for cloud regions and self-hosted instances.

Tracing covers the providers that go through nanoinfra's OpenAI-compatible client path. Native providers that do not use that client may not produce Langfuse OpenAI-wrapper traces.

Metrics and Pricing​

The Metrics destination in the WebUI needs no configuration: Usage, Live and Calls all read what nanoinfra already records. Two things are configurable, and both are off until you set them.

Prices​

The WebUI is the place for this. Open Settings → Models, open a model, open Pricing. That panel has four fields and a live figure: what the recorded window would have cost at the rates you typed. Settings → Providers → Default pricing sets rates for a whole provider. It carries a free checkbox, the one-edit answer for a local one.

The config below is what those panels write, and what to edit when the deployment is configured from a file.

nanoinfra counts tokens and ships no rates, so state yours, per model:

{
"pricing": {
"openai/gpt-4o": {
"inputPerMtok": 2.5,
"outputPerMtok": 10.0,
"cacheReadPerMtok": 1.25,
"cacheWritePerMtok": 3.125
},
"anthropic/claude-sonnet-4-5": {
"inputPerMtok": 3.0,
"outputPerMtok": 15.0,
"cacheReadPerMtok": 0.3,
"cacheWritePerMtok": 3.75
}
}
}
KeyMeaning
pricing["<provider>/<model>"]The model these rates apply to, matching the per-model rows in Metrics → Usage.
inputPerMtokUSD per million prompt tokens.
outputPerMtokUSD per million completion tokens.
cacheReadPerMtokUSD per million tokens read from the prompt cache.
cacheWritePerMtokUSD per million tokens written to it.

Four rates rather than one, because a cached read is not billed like a fresh prompt token. Any rate you leave out is 0. Each key is also accepted in snake case (input_per_mtok).

An entry that sets no rate that nanoinfra recognises — {}, or a misspelled key — reads as unpriced rather than as free. An explicit 0 is a price and means free. A model with no entry shows — in the cost column, and a deployment with no pricing block at all reads "No prices configured".

What it cost in money has the reasoning: why four rates, why zero is a price, and why unpriced is not free.

A default per provider​

A provider whose models all cost the same thing needs one entry, not one per model. In practice that is a local one, where the same thing is nothing:

{
"providers": {
"ollama": {
"pricing": {
"inputPerMtok": 0,
"outputPerMtok": 0,
"cacheReadPerMtok": 0,
"cacheWritePerMtok": 0
}
}
}
}

A model's own entry always wins. With no entry of its own a model takes this default, and with neither it stays unpriced. The Usage table marks a cost that came from a provider default with a *. An inherited figure is a weaker claim than one somebody typed for that model.

Four explicit zeros mean free, which is true for a local model and is why the WebUI offers it as a checkbox. Leaving the block out entirely is not the same thing: that is unpriced, and reads as —.

What happens when a model changes​

Rates are keyed by "<provider>/<model>" and not by the configuration that names it, which matters in three edits:

  • Two configurations naming one model share one bill. The Pricing panel says so by name.
  • Changing a configuration's model carries its rates to the new key and leaves the old entry alone. llm_calls keeps 400 days under the old model. Deleting the entry would un-price rows still on the page.
  • Deleting a configuration leaves its price. Same reason: it is what prices the history.

Clearing rates is explicit — the Clear rates button, or removing the key — because 0 is a valid rate and an empty field therefore cannot mean "remove".

The /metrics endpoint​

The Live tab's gauges are also served in Prometheus text format, on the gateway's own port rather than the WebUI's:

{
"gateway": {
"metricsEnabled": true,
"metricsToken": "a-long-random-string"
}
}
OptionDefaultDescription
gateway.metricsEnabledfalseAnswer GET /metrics on the gateway health port. Off, so no deployment gains a scrape surface by upgrading.
gateway.metricsToken""The bearer token a scrape must present. Empty means only a loopback bind is served.
curl -H "Authorization: Bearer a-long-random-string" http://127.0.0.1:18790/metrics

Two rules decide whether a scrape is answered:

  • With a token set, a scrape must present it as Authorization: Bearer <token>. A wrong one gets 401.
  • With no token, only a loopback bind is served. Any other bind answers 404.

The endpoint also exports counters and a latency histogram, and no series carries a per-turn label. Scraping it has the names, and the reasoning behind both rules.

See Metrics for what each number means and where it is kept.

Gateway Heartbeat​

The gateway can run a protected heartbeat cron job that periodically checks HEARTBEAT.md in the active workspace. This is enabled by default when you run nanoinfra gateway.

{
"gateway": {
"heartbeat": {
"enabled": true,
"intervalS": 1800,
"keepRecentMessages": 8
}
}
}

If HEARTBEAT.md has tasks under ## Active Tasks, the agent executes them and sends only useful/actionable results to the most recently active chat target. If the file has no active tasks, or the result is routine with nothing useful to report, the heartbeat is skipped silently.

This is intentionally different from user-created cron jobs. A cron job created with the cron tool runs as a scheduled turn in its origin chat/session. It delivers the result back to that channel by default.

HEARTBEAT.md is no longer the only way to keep a recurring check quiet. A scheduled automation carries its own delivery policy. It can be set to report only when the result changes, only on failure, or never. See Automations. Heartbeat remains the right choice in two cases. The first is when the checks themselves live in a file you edit rather than in one automation per check. The second is when the decision to speak should come from evaluating the run rather than from comparing it to the last one.

The heartbeat job is backed by the same cron service as user-created reminders. It is stored under the active workspace (<workspace>/cron/jobs.json) and shows up in cron(action="list") as heartbeat. It is system-managed and cannot be removed with the cron tool. Disable it through config and restart the gateway if you do not want periodic heartbeat checks.

OptionDefaultDescription
gateway.heartbeat.enabledtrueRegister the built-in heartbeat cron job on gateway startup.
gateway.heartbeat.intervalS1800Seconds between heartbeat checks.
gateway.heartbeat.keepRecentMessages8Number of recent heartbeat-session messages to retain after each run.
gateway.restartModeautoRestart strategy for /restart: auto uses exec, replacing the current process. Use exit when a service manager such as systemd owns the restart.

Custom heartbeat evaluator prompt​

The notification gate runs on a built-in system prompt. Advanced users can override it, but you rarely need to — it's strongly advised to first read the evaluator code and the default evaluator.md. To override, drop your prompt at <workspace>/prompts/evaluator.md. It must still instruct the model to call the evaluate_notification tool. Otherwise the gate fails closed and stays silent.

Subagent Concurrency​

By default, nanoinfra only allows one spawned subagent at a time. When the limit is reached, the spawn tool returns an error, so the agent can decide to wait or rearrange its work. It does not queue, because a queued subagent is work nobody asked to wait for. The default protects local LLM servers from loading multiple KV caches at once. If your provider can handle more parallel work, raise the limit:

{
"agents": {
"defaults": {
"maxConcurrentSubagents": 2
}
}
}

Subagents also stop immediately when one of their tools returns an execution error. That default keeps failures visible to the parent agent. If your subagent workflows use tools that can fail transiently and should be retried or worked around by the model, disable hard-stop behavior:

{
"agents": {
"defaults": {
"failOnToolError": false
}
}
}
OptionDefaultDescription
agents.defaults.maxConcurrentSubagents1Maximum number of spawned subagents that may run at the same time. Range: 1–8. Attempts to spawn beyond this limit return an error. Editable from Settings → Agents in the WebUI. The gateway reads it at startup, so it asks for a restart.
agents.defaults.failOnToolErrortrueStop a spawned subagent when a tool execution fails. Set to false to return tool errors to the subagent model so it can recover within the same run.

What raising it costs, and what it shares​

Each subagent is a full conversation with the model, so three in parallel is three times the tokens and three times the rate-limit pressure. The number is capped at 8 in the schema rather than only in the UI. So a typo in the config file cannot become a fork bomb against the provider account.

Subagents of the same parent session share one read/write tracker, so read-before-edit and read deduplication work across them. One subagent's read of a file is visible to its sibling. Subagents of different sessions stay isolated, which is the property that tracker exists for. This matters at any limit above 1, because subagents share a workspace. Without it, two of them editing one file would clobber each other, with the very guard meant to catch that looking the other way.

Two other things worth knowing before turning the number up:

  • A subagent cannot spawn another subagent. Not through a depth limit — the spawn tool is absent from the subagent tool registry. There is no nesting to reason about.
  • A subagent turn is unattended for capability gates. Three in parallel are three things nobody is watching. Anything remote needs a standing grant, because no operator can answer an approval from inside one.

Auto Compact​

When a user is idle for longer than a configured threshold, nanoinfra proactively compresses the older part of the session context into a summary. It keeps a recent legal suffix of live messages. This reduces token cost and first-token latency when the user returns. The model receives a compact summary, the most recent live context, and fresh input. It does not re-process a long stale context with an expired KV cache.

{
"agents": {
"defaults": {
"idleCompactAfterMinutes": 15,
"idleCompactCheckIntervalSeconds": 60
}
}
}
OptionDefaultDescription
agents.defaults.idleCompactAfterMinutes15Minutes of idle time before auto-compaction starts. Set to 0 to disable. The default is close to a typical LLM KV cache expiry window, so stale sessions get compacted before the user returns.
agents.defaults.idleCompactCheckIntervalSeconds60Minimum number of seconds between scans for idle sessions. Set to 0 to scan on every idle tick (~1 s).
agents.defaults.midTurnMessages"queue"What happens to a message that arrives while a turn is already running. queue gives it a turn of its own, behind the one in flight — its own turn_id, its own answer. inject folds it into the running turn, which is what every deployment did before this option existed.

A message that arrives mid-turn​

A turn can take minutes and make dozens of provider calls, so a second message often arrives while the first is still being worked on. midTurnMessages decides what happens to it.

queue (the default) gives it a turn of its own. _dispatch takes the session lock, so the message waits for the turn in flight. It then gets its own turn_id, its own answer, and its own record of who asked. This matches the request→response shape of the API: one request, one response.

inject folds it into the running turn instead. A correction reaches work already underway, which is sometimes what you want. In exchange the message has no response of its own, and the turn's turn_id belongs to whoever started it. On a deployment where more than one person shares a conversation, that also means the record attributes the second person's request to the first. gates.identityIndependence reads that attribution.

The trade with queue, stated plainly: a correction no longer reaches a turn already underway — it waits for it. On a turn making 23 calls that is minutes of waiting instead of minutes of being ignored. It is worse to watch and easier to reason about.

Slash commands are unaffected in either mode. /stop is handled before the dispatch lock, and other commands are dispatched inline while a turn runs. A /status that answers only after the work it was asking about is a command nobody can use.

sessionTtlMinutes remains accepted as a legacy alias for backward compatibility, but idleCompactAfterMinutes is the preferred config key going forward.

How it works:

  1. Idle detection: On each idle tick (~1 s), checks whether an idle-session scan is due. By default, the full scan runs at most once per minute.
  2. Background compaction: Idle sessions summarize the older live prefix via LLM and keep the most recent legal suffix (currently 8 messages).
  3. Summary injection: When the user returns, the summary is injected as runtime context (one-shot, not persisted) alongside the retained recent suffix.
  4. Restart-safe resume: The summary is also mirrored into session metadata so it can still be recovered after a process restart.

[!NOTE] Mental model: "summarize older context, keep the freshest live turns, and overwrite the session file with the compact form". It is not a full session.clear(), but it is a write — not a soft cursor move.

Concretely, auto compact rewrites sessions/<key>.jsonl in place. Older messages (including their structured tool_calls / tool_call_id / reasoning_content) are replaced by just the retained recent suffix (currently 8 messages). The archived prefix is preserved only as a plain-text summary appended to memory/history.jsonl, or as a [RAW] ... flattened dump if LLM summarization fails. The original structured JSON of those turns is no longer recoverable from the session file.

This differs from the token-driven soft consolidation that fires when a prompt exceeds the context budget. That path only advances an internal last_consolidated cursor and leaves the session file untouched. The raw tool-call trail stays on disk and can still be replayed or audited. If you rely on that trail for debugging or auditing, set idleCompactAfterMinutes to 0 and let only the token-driven path run.

Unified Session​

By default, each channel × chat ID combination gets its own session. If you use nanoinfra across multiple channels (e.g. Telegram + Discord + CLI) and want them to share the same conversation, enable unifiedSession:

{
"agents": {
"defaults": {
"unifiedSession": true
}
}
}

When enabled, all incoming messages — regardless of which channel they arrive on — are routed into a single shared session. Switching from Telegram to Discord (or any other channel) continues the same conversation.

Behaviorfalse (default)true
Session keychannel:chat_idunified:default
Cross-channel continuityNoYes
/new clearsCurrent channel sessionShared session
/stop finds tasksBy channel sessionBy shared session
Existing session_key_override (e.g. Telegram thread)RespectedStill respected — not overwritten

This is designed for single-user, multi-device setups. It is off by default — existing users see zero behavior change.

Timezone​

Time is context. Context should be precise.

By default, nanoinfra uses UTC for runtime time context. If you want the agent to think in your local time, set agents.defaults.timezone to a valid IANA timezone name:

{
"agents": {
"defaults": {
"timezone": "Asia/Shanghai"
}
}
}

This affects runtime time strings shown to the model, such as runtime context. It also becomes the default timezone for cron schedules when a cron expression omits tz. It is also the default for one-shot at times when the ISO datetime has no explicit offset.

Common examples: UTC, America/New_York, America/Los_Angeles, Europe/London, Europe/Berlin, Asia/Tokyo, Asia/Shanghai, Asia/Singapore, Australia/Sydney.

Need another timezone? Browse the full IANA Time Zone Database.

Skills Marketplace​

skillsMarketplace points the WebUI's Skills browser at a catalogue. One key today:

KeyDefaultMeaning
nanoinfraBaseUrlhttps://skills.nanoinfra.orgBase URL of the nanoinfra skills-server catalogue, used alongside skills.sh as a marketplace provider. Point it at your own deployment to serve an internal catalogue instead.
{
"skillsMarketplace": {
"nanoinfraBaseUrl": "https://skills.internal.example.com"
}
}

The catalogue a deployment browses is a supply chain: a skill is instructions the agent reads and acts on. Pointing this at a host you do not control means trusting whatever that host serves.

Disabled Skills​

nanoinfra ships with built-in skills, and your workspace can also define custom skills under skills/. If you want to hide specific skills from the agent, set agents.defaults.disabledSkills to a list of skill directory names:

{
"agents": {
"defaults": {
"disabledSkills": ["github", "weather"]
}
}
}

Disabled skills are excluded from the main agent's skill summary, from always-on skill injection, and from subagent skill summaries. This is useful when some bundled skills are unnecessary for your deployment or should not be exposed to end users.

OptionDefaultDescription
agents.defaults.disabledSkills[]List of skill directory names to exclude from loading. Applies to both built-in skills and workspace skills.

Tool Hint Max Length​

Tool hints are the short progress messages shown when the agent calls tools (e.g. $ cd …/project && npm test). By default, these are truncated at 40 characters, which can make long commands hard to read.

Set agents.defaults.toolHintMaxLength to control the truncation threshold:

{
"agents": {
"defaults": {
"toolHintMaxLength": 120
}
}
}
OptionDefaultDescription
agents.defaults.toolHintMaxLength40Maximum characters for tool hint display. Range: 20–500. Higher values show more of the command or path. Lower values keep hints compact.