Skip to main content

Troubleshooting

Use this page to isolate where a failure lives. Start with the smallest surface that proves the most: local CLI first, then gateway, then WebUI or chat apps.

Fast Diagnosis Order​

Run these in order:

nanoinfra --version
nanoinfra status
nanoinfra agent -m "Hello!"

Then, only if the CLI works:

nanoinfra gateway

In a source checkout, prefix every command on this page with uv run. The bare name is on your PATH only when nanoinfra is installed as a package. See Install and Quick Start for the four install methods.

uv run nanoinfra status # in a source checkout

This separates failures into layers:

LayerWhat it proves
nanoinfra --versionInstall and shell command discovery
nanoinfra statusConfig path, workspace, environment references, and active provider/model configuration
nanoinfra agent -m "Hello!"Config loading, provider/model access, workspace writes, and agent loop
nanoinfra gatewayChannel startup, cron system jobs, heartbeat, WebUI/WebSocket, and health endpoint

If nanoinfra agent -m "Hello!" fails, fix that before debugging WebUI, Telegram, Discord, Docker, systemd, or any chat app.

If provider/model setup is incomplete, nanoinfra status points to WebUI Settings → Models or the CLI setup wizard, then prints the command to check again.

How to Read nanoinfra status​

nanoinfra status does not call a model. It checks the selected config and workspace, resolves environment references, and validates the local settings required by the active provider/model without constructing a provider client.

The output has this shape:

🐈 nanoinfra Status

Config: /path/to/config.json ✓
Workspace: /path/to/workspace ✓
Model: provider/model-name (preset: primary)
Agent: ✓ provider/model configuration is ready
Provider A: not set
Provider B: ✓
Local Provider: ✓ http://localhost:11434/v1
OAuth Provider: ✓ (OAuth)

Next: nanoinfra agent -m "Hello!"
Status does not call the model or verify network access and credentials.

The last two lines appear only when the Agent line is ready. One row appears per provider family, so a compatibility alias shares the row of the provider it aliases.

Read it like this:

LineGood signWhat to do if it looks wrong
ConfigIt points to the config file you meant to use and shows ✓.Run nanoinfra onboard, or pass --config to nanoinfra agent, gateway, or serve when testing a non-default instance.
WorkspaceIt points to the workspace you meant to use and shows ✓.Run nanoinfra onboard, create the folder, fix permissions, or pass --workspace on commands that support it.
ModelIt shows the active model or the preset name you expect.Set agents.defaults.modelPreset to the intended preset, or check /model if you changed models during a chat session.
AgentIt says provider/model configuration is ready.Follow the printed WebUI or CLI setup route, then run nanoinfra status again.
Provider rowsThe provider used by the active preset shows ✓, an OAuth marker, or a local URL.Configure only the active provider first. It is normal for unused providers to say not set.

If nanoinfra status looks right but nanoinfra agent -m "Hello!" fails, the install and config paths are probably fine. Continue with Provider and Model Problems.

Installation Problems​

Use the same Python command for install checks and module fallback. That may be python or python3.

SymptomCheck
python: command not foundTry python3 --version. Then replace python in docs commands with the command that worked.
curl: command not foundSomething you pasted needs curl to download it, such as uv's own installer. Install curl through your package manager, or install nanoinfra with python -m pip install nanoinfra in a virtual environment, which needs no download step of its own.
Could not download raw.githubusercontent.comYour network, proxy, or firewall blocked the installer script download. Use manual install from PyPI, or configure your proxy and rerun the command.
nanoinfra: command not foundUse the module form, for example python -m nanoinfra ... or python3 -m nanoinfra .... Reinstall with the same Python command, or add that Python's scripts directory to PATH.
No module named nanoinfraYou are running a different Python than the one used for installation. Run python -m pip show nanoinfra or python3 -m pip show nanoinfra, matching the command that installed nanoinfra.
pip is not availableWhen the installer uses a virtual environment, it tries python -m ensurepip --upgrade. If that fails, install pip for that Python, or use a Python installer/distribution that includes pip.
externally-managed-environmentYour system Python blocks global pip installs. Use uv tool install nanoinfra, pipx install nanoinfra, or create a virtual environment. Do not add --break-system-packages for nanoinfra.
Installed under the wrong PythonInstall with uv tool install nanoinfra, which brings its own interpreter, or name the one you want: uv tool install --python 3.12 nanoinfra, or python3.12 -m pip install nanoinfra inside a virtual environment.
Editable source install does not updateFrom the repo root, run python -m pip install -e . again with the Python command used for development, then check python -m nanoinfra --version or nanoinfra --version.
WebUI build tools missingThey are only needed for WebUI development. Packaged installs already include the WebUI bundle.

Config Problems​

Default config path:

~/.nanoinfra/config.json

Default workspace path:

~/.nanoinfra/workspaces/default/

nanoinfra status reads the default config unless you pass explicit paths. Use the same --config and --workspace across status checks and runtime commands when debugging multiple instances:

nanoinfra status --config ./bot-a/config.json --workspace ./bot-a/workspace
nanoinfra agent --config ./bot-a/config.json --workspace ./bot-a/workspace -m "Hello"
nanoinfra gateway --config ./bot-a/config.json --workspace ./bot-a/workspace

Common config mistakes:

SymptomCheck
JSON parse errorValidate commas, braces, and quotes. Most docs examples are partial snippets to merge.
Unknown or missing providerUse provider registry names such as openrouter, anthropic, openai, ollama, vllm, lm_studio, or define a custom OpenAI-compatible provider key under providers and reference that exact key from the active preset.
snake_case vs camelCase confusionBoth are accepted, but docs use camelCase because nanoinfra writes config with aliases such as apiKey, modelPresets, intervalS.
Environment variable error${VAR_NAME} references are resolved at startup. Set the variable before running nanoinfra.
Edited config but behavior did not changeRestart nanoinfra gateway. Long-running processes read config at startup.

After editing config, check the shortest path to an Agent reply:

nanoinfra status

To refresh missing defaults without overwriting existing settings, run:

nanoinfra onboard --refresh

For an interactive choice between resetting and refreshing, run nanoinfra onboard and choose the option that keeps current values and merges missing defaults.

Provider and Model Problems​

First prove the provider in the CLI:

nanoinfra agent -m "Hello!"

Then compare your config against providers.md.

If you need a known-good snippet instead of diagnosis, use provider-cookbook.md.

SymptomLikely cause
401, unauthorized, invalid API keyKey is missing, expired, pasted with whitespace, or under the wrong provider key.
Model not foundThe model ID belongs to a different provider or gateway.
Provider cannot be inferredPin modelPresets.<name>.provider in the active preset instead of using "auto". For legacy direct configs, pin agents.defaults.provider.
Local model connection refusedOllama, vLLM, LM Studio, or another local server is not running, or apiBase points to the wrong port.
Bedrock validation errorCheck AWS region, credentials, model access, model ID, and whether the model supports Converse.
OAuth provider failsRun the matching login command: openai-codex, xai-grok, or github-copilot, normally with --set-main.
Codex OAuth needs a proxySet providers.openaiCodex.proxy before running the login command. The proxy applies to login, token refresh, and Codex API requests.
Codex login runs on a remote/headless machineIn the WebUI, open ChatGPT in your local browser. When the localhost callback page cannot load, copy the full http://localhost:1455/auth/callback?... URL from the address bar and paste it into the WebUI dialog. From the CLI, open the printed URL locally and paste the same callback URL back into the terminal.
Codex login runs in DockerStart the container with docker run -it so the OAuth flow has an interactive terminal.
Codex says a model is not supported with a ChatGPT accountUse provider openai_codex with a Codex model such as openai-codex/gpt-5.6-sol. Do not use the direct-API openai/... prefix with Codex OAuth.
Config says providers.openai_codex conflicts with the built-in providerUnder providers, keep only the canonical openaiCodex settings key and remove a duplicate openai_codex key. A model preset's provider value remains openai_codex.
xAI OAuth needs a proxySet providers.xaiGrok.proxy before login. It applies to OAuth discovery, token exchange/refresh, and Grok subscription requests.
xAI login runs on a remote/headless machineIn the WebUI, finish sign-in in your local browser. If the loopback redirect cannot reach the server, copy the final URL from the address bar into the WebUI dialog. From the CLI, run nanoinfra provider login xai-grok interactively, open the printed URL elsewhere, and paste the final callback URL or authorization code when prompted.
xAI returns 403 or subscription access deniedConfirm the signed-in account has an eligible X Premium / Grok subscription, then run nanoinfra provider login xai-grok again. This provider does not use an xAI API key or X Developer OAuth.
xAI returns 400 invalid-argumentRead the bounded Response body appended to the provider error. Hosted x_search is sent only when xAI's model catalog advertises supportsBackendSearch. The model ID grok-4.5 itself is valid.
xAI model or X Search stops working after an upstream releaseThe integration follows Grok Build's public OAuth/proxy client contract. Update nanoinfra if xAI changes that contract.

Langfuse Problems​

Langfuse tracing is optional and controlled by environment variables.

SymptomCheck
LANGFUSE_SECRET_KEY is set but langfuse is not installedInstall langfuse in the same Python environment that runs nanoinfra, then restart the process.
No traces appearSet LANGFUSE_SECRET_KEY, LANGFUSE_PUBLIC_KEY, and LANGFUSE_BASE_URL before starting nanoinfra.
Wrong Langfuse project or regionCheck that the key pair and LANGFUSE_BASE_URL come from the same Langfuse project/region.
Only some providers traceLangfuse tracing applies to OpenAI-compatible provider calls. Native providers may not use that client path.

See configuration.md#langfuse-observability for setup commands.

Gateway Problems​

nanoinfra gateway is required for WebUI, chat apps, heartbeat, Dream, and long-running channel connections.

Default ports:

SurfaceDefault
Gateway health endpointhttp://127.0.0.1:18790/health
WebUI/WebSocket channelhttp://127.0.0.1:8765
OpenAI-compatible API (nanoinfra serve)http://127.0.0.1:8900

Common gateway checks:

nanoinfra gateway --verbose
SymptomCheck
Port already in useChange gateway.port, channels.websocket.port, or the --port CLI flag for the relevant command.
WebUI opened on 18790 but shows nothing usefulOpen 8765. 18790 is the health endpoint.
Config changes ignoredRestart the gateway.
Startup pauses at Installing optional featureAn enabled channel is missing its Python dependencies. See Slow Optional Channel Dependency Installation.
Heartbeat never runsKeep the gateway running, add tasks under <workspace>/HEARTBEAT.md -> ## Active Tasks, and make sure gateway.heartbeat.enabled is true.
Cron jobs disappeared after switching workspacesCron jobs are workspace-scoped at <workspace>/cron/jobs.json. Check you are using the intended workspace.

Slow Optional Channel Dependency Installation​

Before loading enabled channels, the gateway checks the dependencies declared by their channel manifests. The CLI and WebUI normally install these dependencies when a channel is enabled. Installation during startup is a recovery path for an enabled config whose Python environment no longer has the required packages. That happens, for example, after you edit the config manually, upgrade nanoinfra, or recreate an isolated uv tool/pipx environment. The gateway waits for the install so an enabled channel is not silently skipped. Later starts skip the installation once the dependencies are present.

If access to PyPI is slow in your region, configure pip to use a trusted package index. The installer honors the standard PIP_INDEX_URL environment variable, including when nanoinfra itself was installed with uv tool:

PIP_INDEX_URL=https://your-trusted-mirror.example/simple nanoinfra gateway

For the systemd user service created by nanoinfra gateway install-service, add a drop-in:

systemctl --user edit nanoinfra-gateway.service
[Service]
Environment="PIP_INDEX_URL=https://your-trusted-mirror.example/simple"

Then reload and restart the service:

systemctl --user daemon-reload
systemctl --user restart nanoinfra-gateway.service

For a system-level or custom service, use sudo systemctl edit <unit> instead. Prefer an HTTPS index operated by an organization you trust, and do not put index credentials in commands or logs.

WebUI Problems​

The packaged WebUI is served by the WebSocket channel.

Minimal config:

{
"channels": {
"websocket": {
"enabled": true
}
}
}

Then run:

nanoinfra gateway

Open:

http://127.0.0.1:8765

If accessing from another device, bind the WebSocket channel to 0.0.0.0 and set token or tokenIssueSecret. The WebSocket channel refuses public binds without a token or token issue secret.

See webui.md#lan-access for LAN setup and ../webui/README.md for frontend development.

Chat App Problems​

Before debugging a chat app:

nanoinfra agent -m "Hello!"
nanoinfra channels status
nanoinfra gateway

Then check:

SymptomCheck
Bot never repliesGateway is not running, the channel is not enabled, or the bot/app token is wrong.
Unknown sender ignoredConfigure allowFrom, pairing, or the channel-specific allow list.
Telegram shows a saved configuration but cannot complete a live checkThe token is saved. Confirm the gateway can reach api.telegram.org, or open Settings → Channels → Telegram → Advanced → Network proxy and enter an HTTP or SOCKS proxy.
Telegram rejects the tokenCopy the current token from BotFather or regenerate it.
Telegram receives no messagesConfirm the channel is enabled, the gateway is running, and the sender is paired or listed in allowFrom.
Discord replies missingEnable Message Content intent and invite the bot with the required permissions.
WhatsApp login expiredRe-run nanoinfra channels login whatsapp.
Chat app works but WebUI does notThe provider and gateway are likely fine. Debug the WebSocket channel separately.

See channels.md for channel-specific setup.

Tool and Workspace Problems​

SymptomCheck
File access deniedCheck tools.restrictToWorkspace and whether the target path is inside the active workspace.
Shell commands fail in DockerSandbox settings may need Linux capabilities. See deployment.md.
Web fetch blockedSSRF protection blocks unsafe targets. Use tools.ssrfWhitelist only for trusted private networks.
MCP tools missingCheck tools.mcpServers, server startup command, environment variables, and tool allow list.
Generated artifacts are missingCheck the active workspace and channel media directory.
A remote command, a web tool, or a stdio MCP tool refusesSee Capability Gate Problems.

Capability Gate Problems​

Two kinds of failure look similar and need different fixes. A deployment fault means that a helper process does not answer. A policy decision means that the capability gate refused the action. The wording tells them apart, so read the message before you change config.

For the policy behind these messages, read capability-gates.md. For the three postures and the config that each one needs, read capability-gates.md#choose-a-posture. For the processes, read deployment.md#the-process-split.

Deployment Faults​

The fetcher does not answer. web_fetch and web_search return this:

The fetcher is not reachable: Could not reach the fetcher at <path>: <cause>
This is a deployment fault rather than a web error. Check that the fetcher process
runs. Nothing reached the network, and no other tool answers this request.

There is no fallback path. A fallback fetch would put web egress back beside the credential store. Check the gateway log for gates: the fetcher listens on ..., or for gates: the fetcher did not start. A fetcher never starts while tools.web.enable is false.

The executor does not answer. execute_on_server returns this:

The executor is not reachable, so nothing ran: Could not reach the executor at
<path>: <cause> This is a deployment fault rather than a policy decision. Check
that the executor process is running.

Every gateway start runs an executor, and so does the Python SDK with remote_execution="executor_process". Read the gateway log first:

gates: the executor listens on <path> (pid <pid>)
gates: the executor did not start, so every gated action refuses until an executor
answers: <cause>

A container that started as root runs the executor from its entrypoint instead, and the gateway then starts none. In the container log, look for these lines:

[entrypoint] executor socket ready at /run/nanoinfra-exec/executor.sock
[entrypoint] warning: no executor socket at /run/nanoinfra-exec/executor.sock after 5s
[entrypoint] error: the executor failed 5 starts in a row — no more restarts

A gateway that finds NANOINFRA_EXECUTOR_EXTERNAL in its environment starts no executor of its own, and it says so:

gates: an executor already runs on <path>, so the gateway starts none

Unset that variable when no other supervisor runs an executor.

A helper refuses to start under its confinement. The kernel reported Landlock and then rejected the ruleset. The supervisor names the child:

the executor did not start under its confinement (<cause>). <hint>

The container path exits with status 78 and stops the retry loop at once:

[confinement] error: this helper refuses to start unconfined
[entrypoint] error: the executor refuses to start unconfined
[entrypoint] error: every gated action stays refused

A kernel with no Landlock support is a different case. The helper starts, and the log carries a warning rather than an error. Read deployment.md#per-process-confinement.

An ansible-runner action fails on a projectPath outside the workspace. The executor's confinement grants no write access there, and ansible-runner writes its artifacts under that path. Move the project under the workspace, or under the gateway's working directory.

The approvals inbox reports a problem. Three states differ:

What the screen saysCauseFix
The gateway cannot reach the executorthe operator socket is absent or closedstart the executor, then reload the page
This gateway holds no approvals inboxno WebUI channel runs, so the route answers 503enable the WebSocket channel
No action waits for an answerthe queue is really emptynothing, and policy still refuses what it refuses

A gateway with no WebUI channel logs this at start:

gates: no WebUI channel is enabled, so no operator can answer a suspended action.
Every approve decision waits for gates.approvalTimeoutS and then refuses.

The MCP host does not answer. Every stdio MCP tool fails with this:

Could not reach the MCP host at <path>: <cause>. Stdio MCP servers stay
unavailable until that process runs.

A host that dies mid-session gives a second wording, and the agent can reconnect after it:

The MCP host connection closed for '<server>': <cause>

HTTP and SSE MCP servers do not use the host, so they keep working. The gateway starts no host when your config names no stdio server.

The audit record cannot be written. The executor refuses the action instead of running it:

The executor did not act on '<server>'. The gate decided, and the audit record
could not be written (<cause>). An action that nothing records does not run.

Check the owner and the mode of ~/.nanoinfra/gates. The executor account must own it. A moved or replaced audit directory gives a named cause:

The audit root <path> is not the directory this process opened (pinned device and
inode <pinned>, found <found>). A rename or a replacement of the audit root hides
the records that the denial latches come from, so this store refuses to read it or
to write to it.

Policy Refusals​

A refusal reaches the transcript after the prefix Denied by the capability gate:. The sentences below are the reasons that follow that prefix.

An unattended action has no grant. A cron job, a long-horizon goal, or a subagent reads this:

Refusing mutate.remote at host scope in a unattended context. A standing grant
must list every resolved host (<names>) and the exact command. See
gates.standingGrants.

Write a grant that names every resolved host and the exact command, then set the matching scope to grant. See configuration-security.md#standing-grant-keys.

A grant exists and nothing happens. The refusal names the key to change. Read The Grant Is Right and the Run Is Still Refused for that message and its fix.

An inventory write refuses. create_server, update_server, and delete_server read this on an unattended turn:

mutate.inventory in a unattended context is deny. A standing grant cannot permit
an inventory write.

No grant can lift that. Make the change from a chat session, from the WebUI, or set gates.unattended.mutate.inventory to allow on purpose.

A runtime approval has nobody to ask. A decision of approve refuses when no correct answer can exist. The shipped gates.approvers list is empty, so a chat turn in the WebUI reads this:

gates.approvers lists nobody on a second authenticated path, so nobody may answer
this request. gates.approvalPaths lists ['webui'], and the request arrived on
'websocket'. Add an approver on another path, or declare a standing grant.

Add one entry to gates.approvers. A bare-token WebUI deployment writes { "channel": "webui", "sender": "webui" }. See configuration-security.md#who-may-answer-an-approval.

A deployment whose gates.approvalPaths names the origin channel reads the other refusal:

no second authenticated path is configured. Add one, or declare a standing grant.
gates.approvalPaths lists ['telegram'], and the request arrived on 'telegram'.

Add a second entry to gates.approvalPaths, and add an approver on that path. A recurring action needs a standing grant instead. A human who answers forty prompts a week no longer reads them. See capability-gates.md#path-independence.

Both refusals describe the deployment rather than the action, and both latch the class for that session today. Lift the latch from the WebUI banner. To avoid the latch, choose a posture that needs no runtime approval. See capability-gates.md#choose-a-posture.

An approval answer does not count. The action stays pending, and the screen names the rule. The audit log holds the same event under the decision approval_refused. Read webui.md#how-a-refusal-reads.

The most common cause is a sender that does not match the actor string. The WebUI asserts webui for a bare-token deployment, and webui:<identity> behind a trusted proxy. The comparison is exact. See configuration-security.md#who-may-answer-an-approval.

Nobody answered in time. The action refuses, and the audit log holds one expired record:

no operator answered before the deadline, so the action expired. Ask again when an
approver is present, or declare a standing grant

Raise gates.approvalTimeoutS up to 300, or declare a standing grant. An expiry latches the class for the rest of that session, and a gateway restart does not restore that latch.

approve sits in the unattended block. Nobody waits on that turn:

mutate.remote at group scope is 'approve' for an unattended context, and no person
waits on this turn. A runtime approval there is a hang or a rubber stamp, so set
gates.unattended.mutate.remote.group to 'grant' and declare a standing grant.

The session is latched. After one denial, the next gated action of that class refuses, and nobody is asked. Read The Grant Is Right and the Run Is Still Refused for that message and the mechanism behind it.

Lift the latch from the WebUI banner on that session. Elapsed time never lifts it, a new turn never lifts it, and a gateway restart never lifts it. See webui.md#latch-banner.

A gateway with no WebUI channel logs this at start, and a latched session there has no control that lifts it:

gates: no WebUI channel is enabled, so no operator control can lift a denial
latch. A latched session stays latched until a gateway with the WebUI runs.

The scope will not resolve. Every context refuses, an interactive turn included:

The host set did not resolve, so the blast radius is unknown (<cause>). Add an
inventory the resolver can read, or install ansible-core so the resolver can ask
ansible for its own configuration.

This message reaches you when an ansible-runner server sets inventoryHost to a plain address, and the deployment has no inventory the resolver can read. Add an inventory file under projectPath, or install ansible-core so the resolver can ask ansible for its own configuration.

The target guard refuses earlier in the other cases, and it has its own wording:

Cannot validate network target: <cause>
Cannot validate network target: the server config names no address to check.
Refusing to execute: <error> One blocked host refuses all of them.
CauseFix
The server config names no inventoryHost and no groupSet one of the two fields
The pattern uses syntax the resolver refuses, such as !, &, ~, :, or [Name the hosts or the group plainly
One host of a group is loopback, link-local, or the cloud metadata addressRemove that host from the group, because one blocked host refuses all of them

The pattern is unbounded. The gate refuses before it reads your policy:

The pattern names an unbounded host set, so its scope is `all`. No policy permits
`all` scope, and no approval path exists for it.

Narrow the pattern. A wildcard and the literal all both resolve to scope all, and no config value permits that scope.

The Grant Is Right and the Run Is Still Refused​

Two things outrank a correct grant, and both are easy to miss because the grant itself looks fine in config.

A matrix cell shadows it. A grant never overrules deny. The refusal names the key:

Standing grant <id> covers this action, but
gates.unattended.mutate.remote.host is 'deny'. Set it to 'grant' for the grant to
apply.

The session is latched. A denial latches the capability class for that session. While the latch holds, the gate is not consulted at all. So a grant added afterwards changes nothing, and the refusal says nobody was asked:

Refused by the capability gate: a mutate.remote action was already denied in this
session. Nobody was asked, because that denial stands.

Clear it from the WebUI, which holds the only control that lifts a latch. A restart does not clear it. Latch state is rebuilt from the audit log, precisely so that an agent able to cause a restart cannot clear its own latch.

A commissioning run reports both of these before a schedule ever meets them. If an automation was saved disabled at creation, read its finding first. The finding names the shadowed cell, the credential cell, and the latch when one is in the way. See automations.md#creating-one-rehearses-it.

The command changed by one character. commands matches exact strings, so a grant for uptime refuses uptime -p. A rehearsal observes one run. If the model composes a different command later, that run is refused and the finding on a fresh rehearsal names the new string.

Confirm the Policy in Force​

A mistyped top-level gates key loads as the shipped defaults, because the root config accepts extra keys. Read the line the gateway logs at start rather than the block you believe you wrote:

nanoinfra gateway --verbose
gates: unattended mutate.remote host=deny group=deny all=deny,
mutate.inventory=deny, credential.access=deny, no standing grants, so no
automation may run a remote command (shipped defaults, no gates policy in config).
confinement: landlock abi <abi> on each helper process, filesystem rules, a tcp
port allowlist for the fetcher, and no tcp listener

The words (shipped defaults, no gates policy in config) report one fact: the unattended block holds every shipped value. A mistyped top-level gates key produces those words. A policy that denies every unattended action on purpose produces them too, so compare the values in the line with the values you wrote. A key inside gates that is mistyped fails to load instead, and the error names the key.

The line states the unattended half only. Read the policy panel for the interactive half. See capability-gates.md#confirm-the-posture-is-live.

Every decision also reaches the audit log at ~/.nanoinfra/gates/gate-YYYY-MM-DD.jsonl. Read it in the WebUI under Settings → Security → Gate decisions, or with a shell on the host.

Remote Execution Problems​

SymptomFix
You approved the action and it failed anyway, saying the server needs a credential this process cannot decryptNANOINFRA_SECRETS_KEY is not set, or is not a valid key, in the environment that started the gateway. The secret itself is untouched. Set the variable where the gateway runs and try again.
The message says the server "references secret <id>, which no longer exists"That one is literal: the secret was deleted. Recreate it in Settings → Secrets and point the server at it. Distinguish it from the row above — the two look similar and have different fixes.
An automation refuses to run, naming a referenced server or diagramThe reference no longer resolves, and an automation will not fall back to searching by name. Open the automation and the editor shows the broken reference. Replace or remove it. See Automations.
The action was approved, then nothing appears in the job listA failure before the transport opens records no job. Read ~/.nanoinfra/gates/gate-<date>.jsonl for the decision and the reason that followed it.

The gate audit log is the first place to look for any of these. It records the decision, and a completion record follows it with the outcome. An allow with no completion means the action ended without an exit code this side could read.

Memory and Session Problems​

SymptomCheck
Conversation context seems wrongConfirm the active workspace and session. WebUI chats and chat app threads may use different sessions.
Memory does not update immediatelyDream consolidation is periodic. Recent turns still live in session history.
Old sessions appear after moving configSession files are stored under <workspace>/sessions/. Verify the workspace path.
You want one shared session across devicesSet agents.defaults.unifiedSession intentionally. Otherwise keep separate sessions.
A turn or a Dream run fails naming memory/.dream_cursorThat file holds a cursor and something else is in it. It is refused rather than read as 0, because 0 is a legal value meaning "consolidate everything" — reading a broken file as that would re-consolidate history you already have and stop the history file from ever shrinking. Put the cursor of the last consolidated entry in it, or delete the file to start from nothing.
The history file is far larger than maxHistoryEntriesCompaction runs from the append path, so it trims as turns arrive. A file that stays large means the entries are not consolidated yet, and those are never dropped — check whether Dream is enabled and completing (/dream-log).

Diagram Problems​

SymptomCheck
A save is refused naming a secret fieldThat field holds a reference, not a value. Store the value under Settings → Secrets and put secret://<name> in the field. See Infra Diagrams.
The agent says a change was not saved because no preview was shownThe server records the preview a save is allowed to apply. Ask the agent to preview the change, confirm it, and it applies the same payload.
The agent says the payload is not the one that was previewedYou approved a different change than the one it tried to save. Ask it to preview again.
A gallery row says "changed outside nanoinfra"Something other than the app wrote that file. The content still renders, and its update time is not evidence of anything.
A gallery row says "unreadable"The file is on disk and cannot be parsed. Nothing was deleted. Open it in an editor.
A gallery row says "unusable file name"A diagram id is 32 hex characters. Rename the file, or remove it yourself.
The agent will not move nodes you placedIt cannot, by design. Use Auto layout in the editor.

Collect Useful Evidence​

When opening an issue or asking for help, include:

  • install method and nanoinfra --version.
  • operating system and Python version.
  • the command you ran.
  • relevant nanoinfra status output.
  • sanitized config snippets, especially provider, model, channel, and tool settings.
  • gateway logs from nanoinfra gateway --verbose.
  • whether nanoinfra agent -m "Hello!" works.

Never paste real API keys, bot tokens, OAuth tokens, or private chat IDs into public issues.

If you find a docs mistake, outdated command, or confusing step, please open an issue: https://github.com/nanoinfraorg/nanoinfra/issues.