Secrets and Servers
Store credentials safely, and keep an inventory of the machines you manage. The agent connects to one host and runs something there. The credential never passes through the agent itself.
These are two related modules. Secrets is an encrypted credential store. Servers is an inventory of hosts, and each host record can point at a Secret to authenticate with. Servers also connects to one host and runs a command or an action.
Secrets
Enable it
Secrets needs one environment variable. Set it in the same environment that starts the gateway: a service unit, a container, or a shell profile. nanoinfra never generates or stores this key itself:
export NANOINFRA_SECRETS_KEY="$(python3 -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())")"
Without it, the gateway still starts normally. Every Secrets operation returns a clean "not configured" error until you set the key. Store a copy of this key somewhere safe outside the machine. If you lose it, every secret already stored under it becomes permanently undecryptable, by design. There is no recovery path, and no way to reset it in place.
Create one
There is no agent tool for creating, editing, or rotating a secret — only the WebUI. This is deliberate: the LLM never sees a plaintext credential value, not even for one turn. Open the WebUI's Infrastructure → Secrets page, click New Secret, and fill in:
- Name — must be unique (case-insensitive). This is what a Server's
secretRefand the agent both refer to it by. - Kind —
password,api_key,ssh_key, ortoken. This is a UI hint only (which input widget to show). It does not change how nanoinfra stores the value. - Provider —
local(one encrypted file per secret in the workspace) orpostgres(shared across a deployment viaNANOINFRA_SECRETS_POSTGRES_DSN, see Configuration). Pick per-secret, not instance-wide — a personal SSH key can stay local while a shared production password lives in Postgres. - Value — write-only. Updating a secret always requires typing the full new value. There's no partial edit and the previous value is never shown back to you, by anyone, ever.
No agent tool returns Secrets metadata — not a value, and not even a list of names. If you want the agent to wire a Secret to a Server, tell it the secret's name yourself. It cannot look one up.
Who writes the file
On a deployment with the privilege split, secrets/ belongs to the executor account. The directory carries a group read, so the gateway can list metadata. It carries no group write. That mode is the reason a compromised agent with a shell can enumerate credentials, and cannot replace one. A replacement would swap a refresh token, or repoint a secretRef at a value the agent chose.
So the write crosses the process boundary, and nobody widens the mode. The gateway holds the encryption key and encrypts the value. The executor then writes ciphertext it cannot read. What moves is a file write, not a secret. Three verbs — create, update, delete — and no read. Reading a plaintext has never crossed that wire and does not now.
You will not notice any of this, except in one case. If the executor is not running, a write answers with two facts. The store belongs to another account, and the executor is not reachable. That is a deployment fault rather than a rejected value, and it says so.
A record the executor writes takes the directory's group, not the writer's own. That sounds like an implementation note. It is the difference between a listing that works and one that raises. A directory carrying setgid hands its group down, and one created before that entrypoint does not. A 0640 file in the executor's own primary group is unreadable to the gateway that just asked for it. Fixed in v1.0.6.
Before v1.0.5 this failed outright. The Secrets page could not create a secret on any container deployment, with a raw permission error naming a temporary file. A plain pip install, where one account owns everything, was unaffected — which is why it went unreported.
Where a Stored Value Can Still Appear, and What Removes It
A credential does not stay inside the store. A resolved command embeds one, and mysql -p<password> is the ordinary case rather than the unusual one. So the executor scrubs every stored value out of the text it returns. It does that scrub in its own process, where the plaintext already lives. The agent process decrypts nothing to make this work.
Four parts of a turn take that scrub:
| Part of a turn | What replaces a match |
|---|---|
| the text an operator reads | a placeholder that names the secret |
| a tool argument | a placeholder, whatever the tool declared sensitive |
| the reasoning of the turn | a placeholder, value by value |
the result of a credential.access call | the whole result drops |
The last row differs from the others on purpose. That result is the credential, so nothing in it is worth keeping. Reasoning is the plan of a turn, and an operator reads it to see what the turn did. So one value goes, and the words around it stay.
One visible consequence. A provider signs a reasoning block, and a signature has to match the text it signed. A scrub changes the text, so the stored block loses its signature and carries a marker that says why. A turn whose reasoning held a stored secret therefore replays with no reasoning block. A turn that held no secret keeps its signature. The cost falls on the stored copy, and never on a turn in flight.
A scrub that cannot run keeps nothing. A marker replaces the text instead. A store that fails is a reason to write less, not more.
One session file has five writers, and every one of them scrubs. This is worth stating, because each writer was found by looking for the next one:
What writes into ~/.nanoinfra/sessions/*.jsonl | When |
|---|---|
| the message records | every turn |
| the reasoning of a turn | every turn with a thinking model |
| a checkpoint of a turn still in flight, in the metadata line | every turn that runs a tool |
| the chat-style messages of a provider's replay cache | every save with a cache |
| the provider's own replay payload, on the Responses and Codex paths | every save on those providers |
If you add a persistence path of your own, scrub it. Nothing about the file's shape prevents a sixth writer.
The provider replay cache is the one place where a failed scrub costs something visible. nanoinfra drops the cache rather than writing it. The next turn then replays from the message history, which is slower on the provider's side and loses nothing.
Servers
Create one
Either the agent (create_server, update_server, list_servers, get_server, delete_server — all dry_run-gated, preview first) or the WebUI's Infrastructure → Servers page. Every server has:
[!IMPORTANT] The three write tools carry the capability class
mutate.inventory. An unattended turn cannot use them under the shipped policy, because one inventory write can point a granted name at another address. A chat session can. Seecapability-gates.md.
- Name — unique. This is what the agent and the Diagrams target picker refer to it by.
- Provider —
ssh,ansible-runner,ssm, orapi(below). - Config — provider-specific fields, exact keys below.
- Secret — optional, a dropdown of existing Secrets by name (WebUI) or a Secret's id (agent tool's
secretRef). This is where a Server's credential comes from — inventory CRUD never decrypts it, only stores the reference. - Tags — free-form, for your own grouping/filtering.
Provider config fields
| Provider | Config fields | Notes |
|---|---|---|
ssh | host, port, username | port defaults to 22 if omitted. No username means asyncssh falls back to the local process user — usually not what you want. Set it explicitly. |
ansible-runner | inventoryHost, group, projectPath | Targets inventoryHost if set, else group. A server needs at least one of the two, and a server with neither refuses to execute (nothing concrete to validate or target). projectPath points at the Ansible project directory containing your inventory/playbooks. |
ssm | instanceId, region | AWS Systems Manager Run Command — authenticates via IAM/instance-profile permissions, not a Secret. This provider accepts a secretRef and never uses it. |
api | baseUrl | Calls a specific endpoint on this base URL — never an arbitrary shell command. The agent's command argument for this provider is "<METHOD> <path>" (method optional, defaults to GET). nanoinfra refuses any request that would resolve to a different origin than baseUrl, before it ever sends one. |
Device notes: what a box remembers
Every server has a NOTES.md beside its record — the memory of that machine. Not what happened on
it, which the session transcript already holds, but what stays true about it. This box's disk fills
because of one runaway log. This one needs sudo -n. This one's package manager is held back on
purpose. Without it, every session rediscovers the last incident from whoever remembers it.
<workspace>/servers/<id>.json the record
<workspace>/servers/<id>.NOTES.md its memory
<workspace>/servers/<id>.NOTES.archive.md what rotated out
Keyed by the id, not the name, so renaming a box does not orphan its memory. An automation stores an id and re-reads the name when it fires, for the same reason. Plain markdown, in the workspace, so the file browser reaches it, the backup carries it, and you can edit it by hand.
The one rule an operator has to know
A note does not expire. These are facts about static infrastructure: a conclusion about a box stays true until the box changes. So there is no clock that ages a note out. A note that disagrees with what you see is evidence the infrastructure changed, rather than stale noise. The agent is told to say so and to append what changed, instead of quietly overwriting the old note.
Yours outranks the agent's
An entry carries its author, and the page marks a human's:
## 2026-08-14 09:12 UTC · alberto (operator) · journald is deliberate
The debug level is on purpose. Do not change it.
## 2026-09-03 10:04 UTC · sre-copilot · disk pressure
Vacuumed /var/log/journal from 14G. Per the operator's note above, the debug level stays.
Three consequences, all of them in the code rather than in the prompt:
- The author is never something the writer can claim. An agent's name comes from the turn: the
automation's name, or the agent's. The WebUI stamps your verified identity. Nothing can sign
a note
(operator)except this page. - An agent may revise only its own entries. If it thinks your note is now wrong, it appends an entry saying so. That keeps your words, and it puts the disagreement where you can settle it.
- Your entries never rotate out. The cap exists for what an agent accumulates.
From the WebUI
Infrastructure → Servers → a server shows the panel: newest entry first, operator entries distinguishable at a glance. Add note appends one entry signed as you. Edit hands you the whole markdown file, because this file is yours as much as the agent's.
From the agent
One tool, device_notes, with four actions: read, append, revise_own, read_archive. It is
part of the servers tool group, so an operator who puts that group behind a mention withholds the
notes with it.
The agent does not append on every turn — the contract it is given is append what changes what the
next visitor would do. Do not log routine checks, because four hundred entries of "checked disk,
fine" cost every later reader tokens and bury the three lines that mattered. Past 24,000 characters
the oldest agent entries move to <id>.NOTES.archive.md. Nothing is ever deleted. The archive is
readable from the tool and from the panel.
What a note may not contain
nanoinfra refuses a write that carries credential shapes: -----BEGIN, password=, token=, or a
long hex or base64 run. The refusal names the reason, so the model rewrites the note. Not
silently redacted. Masking a value keeps the sentence honest. Masking a hostname turns a useful note
into a riddle, and nobody notices until somebody acts on it.
The screen applies to the agent, not to you. A person deliberately writing in their own file is not the hazard the refusal exists for, so the WebUI never argues with your edit.
nanoinfra gives the agent two more rules. Nothing enforces them, and they are worth knowing when you read a note. First: no command output pasted wholesale, because a note is a conclusion and the transcript already exists in the session. Second: nothing about who asked, because the device's memory is about the device.
When the agent reads them
Only for a server the turn names, and never in the cached part of the prompt. Mention
@server:db-01 and that box's notes arrive with the turn. A turn that names no server loads none.
An automation declares its references by id, so a scheduled run gets the same thing without anybody
typing an @. And a turn that decides mid-flight that it needs a box's history calls
device_notes with action='read'.
A NOTES.md in every prompt because a server exists would be the knowledge-base mistake at a
smaller scale, and it would grow with your own writing.
Connecting and running something
There is no "run" button in the WebUI — execution only happens through the agent, via execute_on_server. Inventory management (create/edit/list/delete) and execution are deliberately separate surfaces. To run something, just ask in a normal chat:
"Connect to
<server name>and runuptime"
Naming the server instead of describing it
Typing a name means the agent has to find it, which it does by calling list_servers and matching.
That is fine when you can see what it picked. Reference it instead and there is nothing to match:
"Check @server:db-01 for failed units and report anything that is not running"
Type @server: in the composer, and the list narrows as you type. It matches the provider and the
tags as well as the name, so @server:prod finds a host tagged prod. Pick one and the message carries
the readable name while the reference carries the id — so a server renamed later still resolves.
What the reference gives the agent is the id, the name, the provider and the tags. Never the
record. A host, a username and a secretRef are inventory, and reading them is what
get_server is for, under the gate. A mention says this server exists and is called this. The
policy still decides everything after that.
It also brings that box's device notes — what an earlier visit established about it. That is what naming a server buys beyond pinning an identity, and it is the only thing a mention loads content for.
This matters most in an automation, where the agent would otherwise redo the search on every unattended run — see Automations.
The tool is a thin client. It writes one request to a Unix socket and renders the reply. The executor process does the rest:
- It resolves the server, and it refuses an unknown name or an unknown provider.
- It resolves the host set, and it checks every resolved address against the target guard.
- It previews the resolved server, provider, command, and host list when the call asked for a preview. A preview reaches no host and resolves no credential. It reports what a real run would meet: the gate's decision for the current context. It also reports the standing grant that would permit the action when one is missing. Asking is pure, so it costs nothing and authorizes nothing.
- It asks the capability gate when the call asked to execute. The gate reads your policy, the class, the scope, and the execution context.
- It writes one audit record for the decision.
- It suspends the action when the decision is
approve, and it waits for a human answer on a second socket. Seecapability-gates.md#the-approval-path. - It resolves the Secret inside its own process, dials the host, and reports the result.
Two properties follow from that split:
- The agent process holds no transport and no credential value. A test walks the syntax tree of the tool module and fails on an import of a backend or of the credential store.
- The
dry_runargument requests a preview. It does not authorize execution. Policy decides that, and no value on the call changes the answer. What a preview does carry is the answer, so an operator no longer has to reverse-engineer a grant from a refusal.
This is the highest-consequence tool in the system. A refusal is final for that action, and it blocks further remote actions in that session until an operator lifts the block. Read capability-gates.md before you point this at production.
What happens behind the scenes
- Durable job records. nanoinfra writes every execution attempt to disk before it starts:
queued, thenrunning, then a terminal status. A gateway crash mid-run does not lose the record. The next startup reconciles that record tofailed, with an "interrupted by restart" note. A refused action creates no job record and decrypts no credential, because the gate runs before both steps. The gate audit log holds the refusal instead. - Smart timeouts. Not one fixed deadline. The idle clock resets on real activity, such as SSH's streaming output or Ansible Runner's completion signal. An absolute 30-minute ceiling always applies, whatever the activity. For
ansible-runnerandssm, a reported timeout cannot stop the remote work already in flight. Nothing can cancel a thread mid-blocking-call, and nothing can recall an already-sent AWS command. So the agent tells you the command may still be running, rather than pretending otherwise. - Target guard. The guard checks the host address before the executor connects. It blocks loopback, link-local, and the cloud metadata address. It deliberately allows private (RFC1918) ranges, because most real infrastructure lives there. This guard exists so that no Server record can reach the metadata endpoint and exfiltrate cloud credentials. The guard runs in the executor process, and it checks the addresses that the backend will dial.
- Group actions check every host. An
ansible-runnerpattern expands to named hosts first. The guard then checks each one, and one blocked host refuses the whole action. Partial execution never happens. - nanoinfra disables SSH host-key verification (no trust-on-first-use store exists yet). This is an accepted, documented risk (see
.agent/security.md), and not an oversight. Do not point this at a host you do not already trust the network path to.