Configuration
This page is one of four. The reference used to be a single 2,929-line document whose own second section was a hand-written table of contents. It was a page admitting it could not be navigated. It is split by subject now, so the fields for a thing sit beside the pages that explain the thing:
| Page | Holds |
|---|---|
| Configuration | how config is loaded, secrets, environment variables, channels, and the deployment-wide settings |
| Provider and Model Configuration | every provider, model preset, fallback and transcription field |
| Agent and Tool Configuration | named agents, tool groups, web tools, MCP, knowledge and connectors |
| Security Configuration | capability gates, approvers, standing grants and pairing |
Config file: ~/.nanoinfra/config.json
This is the full reference. If this is your first install, start with quick-start.md. If you are trying to choose a model or fix provider/model matching, use providers.md first. Come back here for exact fields and advanced options.
For normal local use, prefer the WebUI before editing JSON:
- Settings → Models manages model choices and provider credentials.
- Settings → Channels guides chat-platform setup.
- Other Settings pages cover built-in capabilities.
- Apps manages CLI App and MCP integrations.
Edit config.json directly when you need an advanced field, automate deployment, or intentionally manage configuration as code.
The JSON examples below are usually partial snippets to merge into your existing config, not full replacement files. For the mental model behind config, workspace, gateway, channels, sessions, tools, and memory, see concepts.md.
The generated config.json uses camelCase keys such as apiKey and intervalS. snake_case keys are also accepted for compatibility, but the docs prefer camelCase because that is what nanoinfra writes back to disk.
For setup and runtime failures, follow the diagnosis order in troubleshooting.md before changing multiple config areas at once.
[!NOTE] If your config file is older than the current schema, run
nanoinfra onboard --refresh. nanoinfra adds missing default fields while preserving your existing values.
Where a Setting Lives
If the WebUI does not expose the option you need, start from the task below. Most advanced changes touch one config section and one verification command.
| Task | First keys to check | Verify with | Deep dive |
|---|---|---|---|
| Make the first model reply work | providers.<name>.apiKey, optional providers.<name>.apiBase, modelPresets.<preset>, agents.defaults.modelPreset | nanoinfra status, then nanoinfra agent -m "Hello!" | Providers, Model Presets |
| Add fallback models | modelPresets.<fallback>, agents.defaults.fallbackModels | nanoinfra status, then a normal agent run | Model Fallbacks |
| Keep secrets out of the config file | ${ENV_VAR} placeholders inside any string value | Start nanoinfra from the same environment that sets the variable | Environment Variables for Secrets |
| Open the bundled WebUI | channels.websocket.enabled, optional channels.websocket.port, channels.websocket.tokenIssueSecret | nanoinfra webui | Channel Settings, WebSocket docs |
| Connect one chat app | channels.<channel>.enabled, channel credentials, optional pairing or channels.<channel>.allowFrom | nanoinfra channels status, then nanoinfra gateway --verbose | Channel Settings, Channels |
| Enable voice transcription | transcription.enabled, transcription.provider, matching providers.<name>.apiKey | Send or upload a short voice message through a configured surface | Transcription Settings |
| Enable web search or fetch | tools.web.search.*, tools.web.fetch.*, optional tools.ssrfWhitelist | Ask a question that requires current web information, then inspect logs if needed | Web Tools, Security |
| Enable image generation | tools.imageGeneration.enabled, tools.imageGeneration.provider, tools.imageGeneration.model, matching provider credentials | Enable Image Generation in the WebUI and send one image request | Image Generation |
| Add external tools through MCP | tools.mcpServers.<name> | Start nanoinfra gateway --verbose and check startup/tool logs | MCP |
| Tighten tool and network safety | tools.restrictToWorkspace, tools.exec.sandbox, tools.ssrfWhitelist, channels.*.allowFrom | Run the same workflow through the channel or CLI you plan to expose | Security, Pairing |
| Let an automation run a remote command again | gates.unattended.mutate.remote, gates.standingGrants | Read the gates: line that nanoinfra gateway logs at start | Capability Gates, Capability gate reference |
| Tune request timeouts or process concurrency | NANOINFRA_LLM_TIMEOUT_S, NANOINFRA_STREAM_IDLE_TIMEOUT_S, NANOINFRA_MAX_CONCURRENT_REQUESTS | Start nanoinfra from the same environment and inspect startup/runtime logs | Runtime Environment Variables |
| Run multiple isolated bots | separate --config and --workspace paths, plus distinct gateway.port or channel ports when processes run together | Use the same explicit paths with nanoinfra status, agent, webui, gateway, and serve | Multiple Instances, CLI Reference |
| Observe model calls | LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, LANGFUSE_BASE_URL environment variables | Run one model call, then check the matching Langfuse project | Langfuse Observability |
| See spend in money, not tokens | pricing["<provider>/<model>"].inputPerMtok and the three other rates | Open Metrics → Usage and check the cost card names a figure rather than "No prices configured" | Metrics and Pricing |
| Scrape the runtime gauges | gateway.metricsEnabled, gateway.metricsToken | curl -H "Authorization: Bearer <token>" http://127.0.0.1:18790/metrics | Metrics and Pricing |
Environment Variables for Secrets
Instead of storing secrets directly in config.json, you can use ${VAR_NAME} references that are resolved from environment variables at startup:
{
"channels": {
"telegram": { "token": "${TELEGRAM_TOKEN}" },
"email": {
"imapPassword": "${IMAP_PASSWORD}",
"smtpPassword": "${SMTP_PASSWORD}"
}
},
"providers": {
"groq": { "apiKey": "${GROQ_API_KEY}" }
}
}
Any string value in config.json can use ${VAR_NAME}. Resolution runs once at startup, in memory only. Resolved values are never written back to disk, so editing config through nanoinfra onboard or the WebUI preserves the placeholder.
If a referenced variable is unset, nanoinfra fails fast and reports the exact config field
and variable name without echoing the field value. Run nanoinfra status with the same
--config path to inspect the problem.
More examples
MCP servers — both stdio env and HTTP headers:
{
"tools": {
"mcpServers": {
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "${GITHUB_TOKEN}" }
},
"remote": {
"url": "https://example.com/mcp/",
"headers": { "Authorization": "Bearer ${REMOTE_MCP_TOKEN}" }
}
}
}
}
Web search providers:
{
"tools": {
"web": {
"search": {
"provider": "brave",
"apiKey": "${BRAVE_API_KEY}"
}
}
}
}
Loading variables at startup
Pick whatever fits your deployment — nanoinfra only reads os.environ at startup, so any mechanism that populates the process environment works.
systemd — use EnvironmentFile= in the service unit to load variables from a file that only the deploying user can read:
# /etc/systemd/system/nanoinfra.service (excerpt)
[Service]
EnvironmentFile=/home/youruser/nanoinfra_secrets.env
User=nanoinfra
ExecStart=...
# /home/youruser/nanoinfra_secrets.env (mode 600, owned by youruser)
TELEGRAM_TOKEN=your-token-here
IMAP_PASSWORD=your-password-here
Docker — pass an env file to the locally built image (one KEY=VALUE per line), or use -e KEY=value:
docker run --rm --env-file=./nanoinfra.env \
-v ~/.nanoinfra:/home/nanoinfra/.nanoinfra \
nanoinfra agent -m "Hello"
direnv — drop a .envrc in your working directory and run direnv allow:
# .envrc (auto-loaded by direnv)
export TELEGRAM_TOKEN=your-token-here
export ANTHROPIC_API_KEY=...
Secret managers (1Password, Bitwarden, pass) — wrap the process so secrets only exist as env vars for the lifetime of the run, never on disk:
# 1Password — references in .env.tpl look like `op://Vault/Item/field`
op run --env-file=.env.tpl -- nanoinfra agent
# pass (passwordstore.org)
ANTHROPIC_API_KEY="$(pass show api/anthropic)" nanoinfra agent
# Bitwarden
ANTHROPIC_API_KEY="$(bw get password api/anthropic)" nanoinfra agent
Runtime Environment Variables
These variables are process-level switches. Set them in the same terminal, service unit, container, or supervisor that starts nanoinfra.
Runtime controls
| Variable | Default | Description |
|---|---|---|
NANOINFRA_MAX_CONCURRENT_REQUESTS | 3 | Maximum concurrently running inbound agent requests. Must be an integer. Set 0 or a negative value for unlimited. |
NANOINFRA_LLM_TIMEOUT_S | 300 | Wall-clock timeout, in seconds. Ordinary requests use this value. Streaming requests use the greater of 300 seconds or twice this value. Set 0 to disable. Sustained-goal turns bypass this wall-clock cap. |
NANOINFRA_STREAM_IDLE_TIMEOUT_S | 90 | Streaming idle timeout, in seconds, used by streaming providers. Invalid or non-positive values are ignored. Values above 3600 are clamped. |
NANOINFRA_OPENAI_COMPAT_TIMEOUT_S | 120 | HTTP request timeout, in seconds, for OpenAI-compatible providers. Invalid or non-positive values are ignored. |
NANOINFRA_WORKSPACE_SANDBOX_ENFORCED | unset | Marks that an external workspace sandbox is already enforced. Truthy values (1, true, yes, on, enabled) use NANOINFRA_WORKSPACE_SANDBOX_PROVIDER as the label. Any other non-false value is treated as the provider name. |
NANOINFRA_WORKSPACE_SANDBOX_PROVIDER | unknown | Display label for the external workspace sandbox when NANOINFRA_WORKSPACE_SANDBOX_ENFORCED is truthy, for example macos_app_sandbox or bwrap. |
NANOINFRA_SANDBOX_ENFORCED | unset | Legacy compatibility alias for NANOINFRA_WORKSPACE_SANDBOX_ENFORCED. |
NANOINFRA_TMUX_SOCKET_DIR | ${TMPDIR:-/tmp}/nanoinfra-tmux-sockets | Socket directory used by the bundled tmux skill scripts. |
Secrets module
| Variable | Default | Description |
|---|---|---|
NANOINFRA_SECRETS_KEY | unset | Master encryption key (Fernet) for the Secrets module. nanoinfra never generates or stores this key — you generate it yourself and set it in the same environment that starts the gateway (service unit, container, or .env): python3 -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())". Without it, the gateway still starts fine. Every Secrets operation just returns a clean "not configured" error until it's set. |
NANOINFRA_SECRETS_POSTGRES_DSN | unset | Postgres connection string for secrets stored with providerId="postgres" (shared across a deployment). Optional — providerId="local" secrets work fully without it. Only Postgres-backed secrets fail with "not configured" if it's unset. Requires the secrets-postgres extra: nanoinfra plugins enable secrets-postgres (or pip install nanoinfra[secrets-postgres]). |
Servers module (execution backends)
No dedicated environment variables — execute_on_server resolves its credential from the Secrets module (above) and connects via whichever provider a Server record specifies. Each remote-connection provider needs its own optional library, all bundled under one servers extra: nanoinfra plugins enable servers (or pip install nanoinfra[servers]) installs asyncssh (the ssh provider), ansible-runner (the ansible-runner provider), and boto3 (the ssm provider). The api provider needs nothing extra — it uses the already-core httpx. A missing library only breaks the one provider that needs it. The gateway starts fine either way, and inventory CRUD (list_servers/get_server/create_server/update_server/delete_server) never needs any of them.
Installer, build, and WebUI development
| Variable | Default | Description |
|---|---|---|
NANOINFRA_BIN_DIR | $HOME/.local/bin | Installer launcher directory on macOS/Linux. |
NANOINFRA_VENV | $HOME/.nanoinfra/venv | Managed virtual environment path used by the installer fallback. |
NANOINFRA_SKIP_WIZARD | unset | Set to 1 to skip the automatic WebUI or wizard setup a first run would start. |
NANOINFRA_SKIP_WEBUI_BUILD | unset | Set to 1 to skip bundling the WebUI during package builds. |
NANOINFRA_FORCE_WEBUI_BUILD | unset | Set to 1 to rebuild the bundled WebUI even when nanoinfra/web/dist/index.html already exists. |
NANOINFRA_EXTRAS | unset | Docker build argument containing comma-separated Python extras such as bedrock. |
NANOINFRA_CHANNELS | whatsapp | Docker build argument containing comma-separated channels whose manifest dependencies are preinstalled. |
NANOINFRA_API_URL | http://127.0.0.1:8765 | Gateway target for the Vite WebUI dev server proxy. |
Internal variables such as NANOINFRA_RESTART_* and NANOINFRA_PATH_* are set by nanoinfra itself and are not a supported user configuration surface.
Channel Settings
Global settings that apply to all channels. Configure under the channels section in ~/.nanoinfra/config.json:
{
"channels": {
"sendProgress": true,
"sendToolHints": true,
"sendMaxRetries": 3,
"telegram": {
"enabled": false
}
}
}
| Setting | Default | Description |
|---|---|---|
sendProgress | true | Stream agent's text progress to the channel |
sendToolHints | true | Stream tool-call hints (e.g. read_file("…")) |
showReasoning | true | Allow channels to surface model reasoning/thinking content (DeepSeek-R1 reasoning_content, Anthropic thinking_blocks, inline <think> tags). Reasoning flows as a dedicated stream with _reasoning_delta / _reasoning_end markers — channels override send_reasoning_delta / send_reasoning_end to render in-place updates. Even with true, channels without those overrides stay no-op silently. Currently surfaced on CLI and WebSocket/WebUI (italic shimmer header, auto-collapses after the stream ends). Telegram / Slack / Discord / Matrix / Mattermost keep the base no-op until their bubble UI is adapted. Independent of sendProgress. |
sendMaxRetries | 3 | Max delivery attempts per outbound message, including the initial send (0-10 configured, minimum 1 actual attempt) |
Non-image attachments are included in the user message as local path references, without
injecting their contents into the model prompt. When file tools are enabled, the agent
can inspect supported text, PDF, DOCX, XLSX, and PPTX files on demand with read_file. It can
also pass the original path to another tool when exact file bytes are required. The deprecated
channels.extractDocumentText setting is accepted for compatibility but ignored.
Normal tool workspace and media access rules still apply to attachment paths.
channels.transcriptionProvider and channels.transcriptionLanguage are deprecated compatibility fields. They remain as a read-only fallback for older configs, but new configuration should use top-level transcription.provider and transcription.language.
sendProgress and sendToolHints can also be overridden per channel. The global values stay as defaults for channels that do not set their own value:
{
"channels": {
"sendProgress": true,
"sendToolHints": true,
"telegram": {
"enabled": true,
"sendProgress": false,
"sendToolHints": false
},
"websocket": {
"enabled": true,
"sendToolHints": true
}
}
}
Telegram richMessages defaults to false. Enable it only to opt in to Bot API 10.1 sendRichMessage rendering. Leave it disabled for Telegram Web clients that show unsupported-message errors for rich messages.
Retry Behavior
Retry is intentionally simple.
When a channel send() raises, nanoinfra retries at the channel-manager layer. By default, channels.sendMaxRetries is 3, and that count includes the initial send.
- Attempt 1: Send immediately.
- Attempt 2: Retry after
1s. - Attempt 3: Retry after
2s. - Higher retry budgets: Backoff continues as
1s,2s,4s, then stays capped at4s. - Transient failures: Network hiccups and temporary API limits often recover on the next attempt.
- Permanent failures: Invalid tokens, revoked access, or banned channels will exhaust the retry budget and fail cleanly.
[!NOTE] This design is deliberate: channel implementations should raise on delivery failure, and the channel manager owns the shared retry policy.
Some channels may still apply small API-specific retries internally. For example, Telegram separately retries timeout and flood-control errors before surfacing a final failure to the manager.
If a channel is completely unreachable, nanoinfra cannot notify the user through that same channel. Watch logs for
Failed to send to {channel} after N attemptsto spot persistent delivery failures.
Langfuse Observability
nanoinfra can trace OpenAI-compatible provider calls through Langfuse's OpenAI SDK wrapper. This is configured with environment variables, not config.json.
Install the optional package in the same Python environment that runs nanoinfra:
nanoinfra plugins enable langfuse
Set Langfuse credentials before starting nanoinfra agent, nanoinfra gateway, or nanoinfra serve:
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
When LANGFUSE_SECRET_KEY is set and the langfuse package is installed, nanoinfra uses langfuse.openai.AsyncOpenAI for OpenAI-compatible providers so model requests are sent to Langfuse in the background. If the secret key is set but langfuse is missing, nanoinfra logs a warning and falls back to the regular OpenAI client.
Use the Langfuse region or self-hosted URL that matches your project. The Langfuse OpenAI SDK docs use LANGFUSE_BASE_URL for cloud regions and self-hosted instances.
Tracing covers the providers that go through nanoinfra's OpenAI-compatible client path. Native providers that do not use that client may not produce Langfuse OpenAI-wrapper traces.
Metrics and Pricing
The Metrics destination in the WebUI needs no configuration: Usage, Live and Calls all read
what nanoinfra already records. Two things are configurable, and both are off until you set them.
Prices
The WebUI is the place for this. Open Settings → Models, open a model, open Pricing. That
panel has four fields and a live figure: what the recorded window would have cost at the rates you
typed. Settings → Providers → Default pricing sets rates for a whole provider. It carries a
free checkbox, the one-edit answer for a local one.
The config below is what those panels write, and what to edit when the deployment is configured from a file.
nanoinfra counts tokens and ships no rates, so state yours, per model:
{
"pricing": {
"openai/gpt-4o": {
"inputPerMtok": 2.5,
"outputPerMtok": 10.0,
"cacheReadPerMtok": 1.25,
"cacheWritePerMtok": 3.125
},
"anthropic/claude-sonnet-4-5": {
"inputPerMtok": 3.0,
"outputPerMtok": 15.0,
"cacheReadPerMtok": 0.3,
"cacheWritePerMtok": 3.75
}
}
}
| Key | Meaning |
|---|---|
pricing["<provider>/<model>"] | The model these rates apply to, matching the per-model rows in Metrics → Usage. |
inputPerMtok | USD per million prompt tokens. |
outputPerMtok | USD per million completion tokens. |
cacheReadPerMtok | USD per million tokens read from the prompt cache. |
cacheWritePerMtok | USD per million tokens written to it. |
Four rates rather than one, because a cached read is not billed like a fresh prompt token. Any
rate you leave out is 0. Each key is also accepted in snake case (input_per_mtok).
An entry that sets no rate that nanoinfra recognises — {}, or a misspelled key — reads as
unpriced rather than as free. An explicit 0 is a price and means free. A model with no entry
shows — in the cost column, and a deployment with no pricing block at all reads "No prices
configured".
What it cost in money has the reasoning: why four rates, why zero is a price, and why unpriced is not free.
A default per provider
A provider whose models all cost the same thing needs one entry, not one per model. In practice that is a local one, where the same thing is nothing:
{
"providers": {
"ollama": {
"pricing": {
"inputPerMtok": 0,
"outputPerMtok": 0,
"cacheReadPerMtok": 0,
"cacheWritePerMtok": 0
}
}
}
}
A model's own entry always wins. With no entry of its own a model takes this default, and with
neither it stays unpriced. The Usage table marks a cost that came from a provider default with a
*. An inherited figure is a weaker claim than one somebody typed for that model.
Four explicit zeros mean free, which is true for a local model and is why the WebUI offers it
as a checkbox. Leaving the block out entirely is not the same thing: that is unpriced, and reads
as —.
What happens when a model changes
Rates are keyed by "<provider>/<model>" and not by the configuration that names it, which
matters in three edits:
- Two configurations naming one model share one bill. The Pricing panel says so by name.
- Changing a configuration's model carries its rates to the new key and leaves the old entry
alone.
llm_callskeeps 400 days under the old model. Deleting the entry would un-price rows still on the page. - Deleting a configuration leaves its price. Same reason: it is what prices the history.
Clearing rates is explicit — the Clear rates button, or removing the key — because 0 is a
valid rate and an empty field therefore cannot mean "remove".
The /metrics endpoint
The Live tab's gauges are also served in Prometheus text format, on the gateway's own port rather than the WebUI's:
{
"gateway": {
"metricsEnabled": true,
"metricsToken": "a-long-random-string"
}
}
| Option | Default | Description |
|---|---|---|
gateway.metricsEnabled | false | Answer GET /metrics on the gateway health port. Off, so no deployment gains a scrape surface by upgrading. |
gateway.metricsToken | "" | The bearer token a scrape must present. Empty means only a loopback bind is served. |
curl -H "Authorization: Bearer a-long-random-string" http://127.0.0.1:18790/metrics
Two rules decide whether a scrape is answered:
- With a token set, a scrape must present it as
Authorization: Bearer <token>. A wrong one gets401. - With no token, only a loopback bind is served. Any other bind answers
404.
The endpoint also exports counters and a latency histogram, and no series carries a per-turn label. Scraping it has the names, and the reasoning behind both rules.
See Metrics for what each number means and where it is kept.
Gateway Heartbeat
The gateway can run a protected heartbeat cron job that periodically checks HEARTBEAT.md in the active workspace. This is enabled by default when you run nanoinfra gateway.
{
"gateway": {
"heartbeat": {
"enabled": true,
"intervalS": 1800,
"keepRecentMessages": 8
}
}
}
If HEARTBEAT.md has tasks under ## Active Tasks, the agent executes them and sends only useful/actionable results to the most recently active chat target. If the file has no active tasks, or the result is routine with nothing useful to report, the heartbeat is skipped silently.
This is intentionally different from user-created cron jobs. A cron job created with the cron tool runs as a scheduled turn in its origin chat/session. It delivers the result back to that channel by default.
HEARTBEAT.md is no longer the only way to keep a recurring check quiet. A scheduled automation carries its own delivery policy. It can be set to report only when the result changes, only on failure, or never. See Automations. Heartbeat remains the right choice in two cases. The first is when the checks themselves live in a file you edit rather than in one automation per check. The second is when the decision to speak should come from evaluating the run rather than from comparing it to the last one.
The heartbeat job is backed by the same cron service as user-created reminders. It is stored under the active workspace (<workspace>/cron/jobs.json) and shows up in cron(action="list") as heartbeat. It is system-managed and cannot be removed with the cron tool. Disable it through config and restart the gateway if you do not want periodic heartbeat checks.
| Option | Default | Description |
|---|---|---|
gateway.heartbeat.enabled | true | Register the built-in heartbeat cron job on gateway startup. |
gateway.heartbeat.intervalS | 1800 | Seconds between heartbeat checks. |
gateway.heartbeat.keepRecentMessages | 8 | Number of recent heartbeat-session messages to retain after each run. |
gateway.restartMode | auto | Restart strategy for /restart: auto uses exec, replacing the current process. Use exit when a service manager such as systemd owns the restart. |
Custom heartbeat evaluator prompt
The notification gate runs on a built-in system prompt. Advanced users can override it, but you rarely need to — it's strongly advised to first read the evaluator code and the default evaluator.md. To override, drop your prompt at <workspace>/prompts/evaluator.md. It must still instruct the model to call the evaluate_notification tool. Otherwise the gate fails closed and stays silent.
Subagent Concurrency
By default, nanoinfra only allows one spawned subagent at a time. When the limit is reached, the spawn tool returns an error, so the agent can decide to wait or rearrange its work. It does not queue, because a queued subagent is work nobody asked to wait for. The default protects local LLM servers from loading multiple KV caches at once. If your provider can handle more parallel work, raise the limit:
{
"agents": {
"defaults": {
"maxConcurrentSubagents": 2
}
}
}
Subagents also stop immediately when one of their tools returns an execution error. That default keeps failures visible to the parent agent. If your subagent workflows use tools that can fail transiently and should be retried or worked around by the model, disable hard-stop behavior:
{
"agents": {
"defaults": {
"failOnToolError": false
}
}
}
| Option | Default | Description |
|---|---|---|
agents.defaults.maxConcurrentSubagents | 1 | Maximum number of spawned subagents that may run at the same time. Range: 1–8. Attempts to spawn beyond this limit return an error. Editable from Settings → Agents in the WebUI. The gateway reads it at startup, so it asks for a restart. |
agents.defaults.failOnToolError | true | Stop a spawned subagent when a tool execution fails. Set to false to return tool errors to the subagent model so it can recover within the same run. |
What raising it costs, and what it shares
Each subagent is a full conversation with the model, so three in parallel is three times the tokens and three times the rate-limit pressure. The number is capped at 8 in the schema rather than only in the UI. So a typo in the config file cannot become a fork bomb against the provider account.
Subagents of the same parent session share one read/write tracker, so read-before-edit and read deduplication work across them. One subagent's read of a file is visible to its sibling. Subagents of different sessions stay isolated, which is the property that tracker exists for. This matters at any limit above 1, because subagents share a workspace. Without it, two of them editing one file would clobber each other, with the very guard meant to catch that looking the other way.
Two other things worth knowing before turning the number up:
- A subagent cannot spawn another subagent. Not through a depth limit — the
spawntool is absent from the subagent tool registry. There is no nesting to reason about. - A subagent turn is unattended for capability gates. Three in parallel are three things nobody is watching. Anything remote needs a standing grant, because no operator can answer an approval from inside one.
Auto Compact
When a user is idle for longer than a configured threshold, nanoinfra proactively compresses the older part of the session context into a summary. It keeps a recent legal suffix of live messages. This reduces token cost and first-token latency when the user returns. The model receives a compact summary, the most recent live context, and fresh input. It does not re-process a long stale context with an expired KV cache.
{
"agents": {
"defaults": {
"idleCompactAfterMinutes": 15,
"idleCompactCheckIntervalSeconds": 60
}
}
}
| Option | Default | Description |
|---|---|---|
agents.defaults.idleCompactAfterMinutes | 15 | Minutes of idle time before auto-compaction starts. Set to 0 to disable. The default is close to a typical LLM KV cache expiry window, so stale sessions get compacted before the user returns. |
agents.defaults.idleCompactCheckIntervalSeconds | 60 | Minimum number of seconds between scans for idle sessions. Set to 0 to scan on every idle tick (~1 s). |
agents.defaults.midTurnMessages | "queue" | What happens to a message that arrives while a turn is already running. queue gives it a turn of its own, behind the one in flight — its own turn_id, its own answer. inject folds it into the running turn, which is what every deployment did before this option existed. |
A message that arrives mid-turn
A turn can take minutes and make dozens of provider calls, so a second message often arrives while
the first is still being worked on. midTurnMessages decides what happens to it.
queue (the default) gives it a turn of its own. _dispatch takes the session lock, so the
message waits for the turn in flight. It then gets its own turn_id, its own answer, and its own
record of who asked. This matches the request→response shape of the API: one request, one response.
inject folds it into the running turn instead. A correction reaches work already underway,
which is sometimes what you want. In exchange the message has no response of its own, and the
turn's turn_id belongs to whoever started it. On a deployment where more than one person shares a
conversation, that also means the record attributes the second person's request to the first.
gates.identityIndependence reads that attribution.
The trade with queue, stated plainly: a correction no longer reaches a turn already underway — it
waits for it. On a turn making 23 calls that is minutes of waiting instead of minutes of being
ignored. It is worse to watch and easier to reason about.
Slash commands are unaffected in either mode. /stop is handled before the dispatch lock, and other
commands are dispatched inline while a turn runs. A /status that answers only after the
work it was asking about is a command nobody can use.
sessionTtlMinutes remains accepted as a legacy alias for backward compatibility, but idleCompactAfterMinutes is the preferred config key going forward.
How it works:
- Idle detection: On each idle tick (~1 s), checks whether an idle-session scan is due. By default, the full scan runs at most once per minute.
- Background compaction: Idle sessions summarize the older live prefix via LLM and keep the most recent legal suffix (currently 8 messages).
- Summary injection: When the user returns, the summary is injected as runtime context (one-shot, not persisted) alongside the retained recent suffix.
- Restart-safe resume: The summary is also mirrored into session metadata so it can still be recovered after a process restart.
[!NOTE] Mental model: "summarize older context, keep the freshest live turns, and overwrite the session file with the compact form". It is not a full
session.clear(), but it is a write — not a soft cursor move.Concretely, auto compact rewrites
sessions/<key>.jsonlin place. Older messages (including their structuredtool_calls/tool_call_id/reasoning_content) are replaced by just the retained recent suffix (currently 8 messages). The archived prefix is preserved only as a plain-text summary appended tomemory/history.jsonl, or as a[RAW] ...flattened dump if LLM summarization fails. The original structured JSON of those turns is no longer recoverable from the session file.This differs from the token-driven soft consolidation that fires when a prompt exceeds the context budget. That path only advances an internal
last_consolidatedcursor and leaves the session file untouched. The raw tool-call trail stays on disk and can still be replayed or audited. If you rely on that trail for debugging or auditing, setidleCompactAfterMinutesto0and let only the token-driven path run.
Unified Session
By default, each channel × chat ID combination gets its own session. If you use nanoinfra across multiple channels (e.g. Telegram + Discord + CLI) and want them to share the same conversation, enable unifiedSession:
{
"agents": {
"defaults": {
"unifiedSession": true
}
}
}
When enabled, all incoming messages — regardless of which channel they arrive on — are routed into a single shared session. Switching from Telegram to Discord (or any other channel) continues the same conversation.
| Behavior | false (default) | true |
|---|---|---|
| Session key | channel:chat_id | unified:default |
| Cross-channel continuity | No | Yes |
/new clears | Current channel session | Shared session |
/stop finds tasks | By channel session | By shared session |
Existing session_key_override (e.g. Telegram thread) | Respected | Still respected — not overwritten |
This is designed for single-user, multi-device setups. It is off by default — existing users see zero behavior change.
Timezone
Time is context. Context should be precise.
By default, nanoinfra uses UTC for runtime time context. If you want the agent to think in your local time, set agents.defaults.timezone to a valid IANA timezone name:
{
"agents": {
"defaults": {
"timezone": "Asia/Shanghai"
}
}
}
This affects runtime time strings shown to the model, such as runtime context. It also becomes the default timezone for cron schedules when a cron expression omits tz. It is also the default for one-shot at times when the ISO datetime has no explicit offset.
Common examples: UTC, America/New_York, America/Los_Angeles, Europe/London, Europe/Berlin, Asia/Tokyo, Asia/Shanghai, Asia/Singapore, Australia/Sydney.
Need another timezone? Browse the full IANA Time Zone Database.
Skills Marketplace
skillsMarketplace points the WebUI's Skills browser at a catalogue. One key today:
| Key | Default | Meaning |
|---|---|---|
nanoinfraBaseUrl | https://skills.nanoinfra.org | Base URL of the nanoinfra skills-server catalogue, used alongside skills.sh as a marketplace provider. Point it at your own deployment to serve an internal catalogue instead. |
{
"skillsMarketplace": {
"nanoinfraBaseUrl": "https://skills.internal.example.com"
}
}
The catalogue a deployment browses is a supply chain: a skill is instructions the agent reads and acts on. Pointing this at a host you do not control means trusting whatever that host serves.
Disabled Skills
nanoinfra ships with built-in skills, and your workspace can also define custom skills under skills/. If you want to hide specific skills from the agent, set agents.defaults.disabledSkills to a list of skill directory names:
{
"agents": {
"defaults": {
"disabledSkills": ["github", "weather"]
}
}
}
Disabled skills are excluded from the main agent's skill summary, from always-on skill injection, and from subagent skill summaries. This is useful when some bundled skills are unnecessary for your deployment or should not be exposed to end users.
| Option | Default | Description |
|---|---|---|
agents.defaults.disabledSkills | [] | List of skill directory names to exclude from loading. Applies to both built-in skills and workspace skills. |
Tool Hint Max Length
Tool hints are the short progress messages shown when the agent calls tools (e.g. $ cd …/project && npm test). By default, these are truncated at 40 characters, which can make long commands hard to read.
Set agents.defaults.toolHintMaxLength to control the truncation threshold:
{
"agents": {
"defaults": {
"toolHintMaxLength": 120
}
}
}
| Option | Default | Description |
|---|---|---|
agents.defaults.toolHintMaxLength | 40 | Maximum characters for tool hint display. Range: 20–500. Higher values show more of the command or path. Lower values keep hints compact. |