Skip to main content

Architecture

This page maps nanoinfra's runtime behavior to source files. Use it when you are debugging internals, reviewing a PR, adding a provider/channel/tool, or trying to understand where a user-visible behavior comes from.

For the product-level mental model, read concepts.md first.

Core Flow​

Main files:

AreaFiles
Message events and queuenanoinfra/bus/events.py, nanoinfra/bus/queue.py
Turn orchestrationnanoinfra/agent/loop.py
Provider/tool conversation loopnanoinfra/agent/runner.py
Context constructionnanoinfra/agent/context.py
Session storage and compactionnanoinfra/session/manager.py
Long-term memory and Dreamnanoinfra/agent/memory.py
Slash command routingnanoinfra/command/
Token and cost accountingnanoinfra/llm_usage/
Credential storenanoinfra/secrets/
DM sender approvalnanoinfra/pairing/
Python SDK entry pointsnanoinfra/sdk/, nanoinfra/nanoinfra.py
OpenAI-compatible HTTP APInanoinfra/api/server.py

Agent Loop vs Agent Runner​

AgentLoop owns the channel-facing turn:

  • receives inbound messages.
  • determines the effective session and workspace scope.
  • builds context.
  • wires hooks, progress, and channel metadata.
  • publishes outbound messages.

AgentRunner owns the model-facing loop:

  • sends messages to the selected provider.
  • handles streaming deltas and reasoning blocks.
  • executes tool calls.
  • feeds tool results back into the model.
  • stops when a final answer is produced or runtime limits are hit.

Keep this split in mind when debugging. If a problem is about channel routing, session keys, workspace selection, or outbound delivery, start in agent/loop.py. If it is about provider calls, tool calls, streaming, or iteration limits, start in agent/runner.py.

Which Agent Answers​

A deployment can configure named agents beside agents.defaults, and each one inherits the defaults and narrows them. AgentLoop._acting_agent_for resolves the acting agent once per turn, from two sources with the explicit one winning:

OrderSourceSet by
1metadata["agent"]the WebUI composer's picker, or an automation's binding
2@agent:<name> in the message texta person, mid-sentence
3nothingfalls through to agents.defaults

A name is a request. nanoinfra validates it against the roster in config. An unknown name falls back to the deployment default rather than raising. So a client cannot name authority into existence. Resolution happens once, before any tool asks what it may reach. So two records cannot disagree about who answered: the request context a tool reads, and the record the transcript keeps.

Main files: nanoinfra/agent/loop.py for resolution, nanoinfra/config/schema.py for the roster and its validators, nanoinfra/cron/agent_binding.py for an automation's bound agent.

Providers​

Provider metadata is centralized in nanoinfra/providers/registry.py. Configuration fields live in nanoinfra/config/schema.py.

Provider selection uses:

  • explicit agents.defaults.provider or preset provider.
  • provider registry keywords.
  • API key prefixes and API base URL hints.
  • local provider fallback when apiBase is configured.
  • gateway fallback for providers that can route many model families.

Provider implementations live in nanoinfra/providers/. Most hosted providers use the OpenAI-compatible implementation, while Anthropic, Azure OpenAI, AWS Bedrock, OpenAI Codex, and GitHub Copilot have specialized paths.

Useful docs:

Channels​

Channels translate external platforms into InboundMessage events and send OutboundMessage events back to the platform.

Main files:

AreaFiles
Base channel contractnanoinfra/channels/base.py
Channel packagesnanoinfra/channels/<channel>/
Discovery and lifecyclenanoinfra/channels/manager.py
WebSocket/WebUI channelnanoinfra/channels/websocket/
Channel access controlchannel config in nanoinfra/channels/*.py

Channels are discovered by scanning self-contained packages under nanoinfra/channels/. Add a channel by contributing one package that follows channel-package-guide.md.

WebUI and Gateway​

nanoinfra gateway starts:

  • enabled chat channels.
  • the WebSocket channel when configured.
  • workspace-scoped cron service.
  • system jobs such as Dream and heartbeat.
  • the health endpoint on gateway.port.

The packaged WebUI is served by the WebSocket channel, not the health endpoint:

SurfaceDefault
Health endpointhttp://127.0.0.1:18790/health
WebUI/WebSockethttp://127.0.0.1:8765

WebUI source lives in webui/. The production build is written to nanoinfra/web/dist/ and bundled into the wheel.

Useful docs:

Process Split​

A running deployment holds five processes: the agent and four confined helpers. Each helper listens on at least one Unix socket, and one request is written per connection.

The agent reaches three of the helpers. The executor reaches the connector host, one hop further out. Its socket group deliberately excludes the agent. A connector call begins in the executor, after the gate has answered. So nothing in the process the model steers has a reason to reach it.

The executor answers the agent on two sockets. executor.sock carries execute requests, whose fields are structured — deliberately, so a caller cannot smuggle free-form model text into an approval prompt. A redaction request is nothing but free-form model text, so it gets its own wire. Two more reasons follow. An execute connection blocks for as long as an operator takes to answer an approval. Persisting a transcript must not queue behind that wait. And the two frames share no field, so one decoder would have to accept two shapes.

The scrub exists because the credential store lives behind the executor. Building redaction sentinels in the agent meant decrypting every workspace secret inside the process the model runs in, on every turn that persisted a transcript. Now the agent sends text and the side that owns the store does the scrubbing. A slow or unreachable scrubber costs a turn its transcript text. A marker is persisted instead. It never costs a raw credential in a durable file.

The transport is a Unix socket in every case. A TCP socket would widen an egress policy to include the executor itself, which holes the one rule that makes a narrow network policy useful.

SocketDefault pathBound by
Execute~/.nanoinfra/run/executor.sockthe executor
Redaction (scrub)~/.nanoinfra/run/executor.scrub.sockthe executor
Operator answers~/.nanoinfra/run/operator/executor.op.sockthe executor
Fetch and search~/.nanoinfra/run/fetcher.sockthe fetcher
Stdio MCP~/.nanoinfra/run/mcp_host.sockthe MCP host
Connector calls~/.nanoinfra/run/connector_host.sockthe connector host

In a container each helper gets its own directory under /run, named for its account: /run/nanoinfra-exec, /run/nanoinfra-fetch, /run/nanoinfra-mcp, /run/nanoinfra-connector. One directory per account rather than one shared one. The reason: write rights on a parent directory allow renaming any entry inside it. So a directory a caller can write is a directory where a caller could substitute its own socket. See deployment.md#the-process-split.

Main files:

AreaFiles
Helper packagesnanoinfra/gates/executor/, nanoinfra/gates/fetcher/, nanoinfra/gates/mcp_host/, nanoinfra/gates/connector_host/
Agent-side thin clientsnanoinfra/agent/tools/server_execution.py, nanoinfra/gates/executor/client.py, nanoinfra/gates/fetcher/client.py, nanoinfra/gates/mcp_host/client.py
Helper serversnanoinfra/gates/executor/server.py, nanoinfra/gates/fetcher/server.py, nanoinfra/gates/mcp_host/server.py, nanoinfra/gates/connector_host/server.py
Connector call, executor sidenanoinfra/gates/executor/connector_action.py, nanoinfra/gates/connector_host/client.py
Wire formatsnanoinfra/gates/*/protocol.py
Child process controlnanoinfra/gates/*/supervisor.py
Process entry pointsnanoinfra/gates/*/__main__.py
Operator answersnanoinfra/gates/executor/operator_socket.py

Import direction is the security property, not a style choice. An agent-side client imports no backend, no SecretStore, and no helper server module. Tests walk the whole syntax tree of each client and fail on such an import, so a lazy import inside a function does not pass.

The same rule covers the approval answer. The approvals inbox runs inside the gateway process, so a second import closure keeps every tool module away from nanoinfra/webui/approvals_api.py.

Per-Process Confinement​

Each helper child starts under a Landlock ruleset. nanoinfra/gates/confinement.py builds one plan per role and applies it in the child, after the fork and before the exec.

PieceWhere
Rules, probe, and per-role plansnanoinfra/gates/confinement.py
Rolesexecutor, fetcher, mcp-host, connector-host
Plan and spawn wiringnanoinfra/gates/*/supervisor.py
Container launcherpython -m nanoinfra.gates.confinement --role <role>
Startup echo clausenanoinfra/gates/startup.py

Three design points shape the module:

  • It calls the kernel through ctypes, so no new third-party package joins the install.
  • It reads a role and never an argv, so a caller of the launcher cannot choose a program.
  • It builds the plan as a value, so a test reads the whole policy without a kernel.

The connector host's plan is the one that differs on purpose. It grants read on connectors/ and not on the workspace, because the workspace holds the credential store. Moving the call out of the executor is only worth anything if the process making a stranger's request holds no secrets.

A kernel with no Landlock support degrades with a warning. A kernel that reports an ABI and then rejects the ruleset refuses the start. For the policy of each role, the failure modes, and the stated limits, read deployment.md#per-process-confinement.

Capability Gate Internals​

The executor holds the gate, because the executor is the process that opens a transport.

AreaFiles
Policy schema and defaultsnanoinfra/config/gates.py
Policy evaluationnanoinfra/gates/policy.py
Capability classesnanoinfra/agent/tools/capabilities.py
Execution contextnanoinfra/agent/tools/context.py
Scope resolutionnanoinfra/servers/scope.py
Network target guardnanoinfra/servers/network_guard.py
Append-only audit lognanoinfra/gates/audit.py
Denial latchnanoinfra/gates/latch.py, nanoinfra/gates/latch_restore.py
Boot-time wiringnanoinfra/gates/runtime.py, nanoinfra/gates/startup.py, nanoinfra/cli/gateway_runtime.py
Approval tokens and payloadnanoinfra/gates/tokens.py, nanoinfra/gates/prompt.py, nanoinfra/gates/approvals.py, nanoinfra/gates/pending.py
Per-process confinementnanoinfra/gates/confinement.py
Operator surfacesnanoinfra/webui/approvals_api.py, nanoinfra/webui/latch_api.py, nanoinfra/webui/audit_api.py, nanoinfra/webui/settings_api.py

The gateway builds the gate runtime once at boot. It keeps the latch controller on the operator side and passes only the gate half toward the tools. So no tool path can clear a latch.

The gateway also starts the helper children before it builds the tools, because each tool resolves its socket path at that moment. The executor starts first and stops last, because the approvals inbox derives its own socket from the path that start exports.

Useful docs:

Tools​

Tools are discovered from nanoinfra/agent/tools/ and plugin entry points.

Important files:

Tool areaFiles
Tool base and schemananoinfra/agent/tools/base.py, nanoinfra/agent/tools/schema.py
Discoverynanoinfra/agent/tools/registry.py
Shell executionnanoinfra/agent/tools/shell.py
Filesystem toolsnanoinfra/agent/tools/filesystem.py
Web search/fetch, with the SSRF/network checksnanoinfra/agent/tools/web.py, nanoinfra/security/network.py
MCP toolsnanoinfra/agent/tools/mcp.py
Cronnanoinfra/agent/tools/cron.py, nanoinfra/cron/
Triggers, the durable file-drop queuenanoinfra/triggers/
Data connectorsnanoinfra/connectors/, nanoinfra/connectors/credentials.py, nanoinfra/gates/connector_host/
Knowledge base and citationsnanoinfra/knowledge/
Infra diagramsnanoinfra/diagrams/
CLI appsnanoinfra/apps/
Audio transcriptionnanoinfra/audio/
Automation statenanoinfra/agent/tools/automation_state.py, nanoinfra/automations/
Commissioningnanoinfra/automations/commissioning.py, commissioning_runner.py, commissioning_state.py, nanoinfra/webui/commissioning_api.py
Image generationnanoinfra/agent/tools/image_generation.py
Runtime self-inspectionnanoinfra/agent/tools/self.py

Tool behavior is part of the model contract. Keep user-visible tool names, schemas, and error messages stable unless a change is intentional.

Config and Paths​

The config schema lives in nanoinfra/config/schema.py. Loading and saving live in nanoinfra/config/loader.py. Runtime path helpers live in nanoinfra/config/paths.py.

Defaults:

PathDefault
Config~/.nanoinfra/config.json
Workspace~/.nanoinfra/workspaces/default/
Sessions<workspace>/sessions/*.jsonl
Memory<workspace>/memory/
Cron store<workspace>/cron/jobs.json
WebUI/media/log runtime dataconfig directory subdirectories such as webui/, media/, and logs/

The schema accepts both camelCase and snake_case keys, but saves config with camelCase aliases.

Agent-Owned State vs Effective Project Context​

Runtime code distinguishes the configured agent workspace from the effective project workspace carried by a session scope. They are often the same path, but a WebUI chat may select a separate project:

ConcernPath owner
Sessions, SOUL.md, USER.md, memory, and custom skillsConfigured agent workspace
Project AGENTS.md, relative tool paths, and shell working directoryEffective project workspace
Workspace access mode and project metadataSession workspace scope

ContextBuilder combines project instructions with agent-owned profile and memory. Filesystem and search tools use the project as their ordinary boundary. They receive only capability-specific read access to built-in/agent skills and the exact agent history file. Keep those cross-root capabilities read-only and explicit. Do not treat the entire agent workspace as an allowed root.

nanoinfra/security/workspace_access.py and nanoinfra/security/workspace_policy.py hold the workspace scope checks.

Memory and Sessions​

Session history is the near-term conversation replay. Memory is the longer-term workspace state.

StoreFile area
Session JSONL files<workspace>/sessions/
Long-term memory<workspace>/memory/MEMORY.md
Consolidation source history<workspace>/memory/history.jsonl
Bootstrap identity files<workspace>/SOUL.md, <workspace>/USER.md, templates under nanoinfra/templates/

Dream is implemented in nanoinfra/agent/memory.py and scheduled by the runtime when enabled.

Security Boundaries​

These code paths are the security-sensitive ones. The section that owns each boundary carries its files, so one path is documented in one place:

BoundaryOwning section
Workspace scopeAgent-Owned State vs Effective Project Context
Shell sandboxingTools
SSRF/network checksTools
Capability gate policyCapability Gate Internals
Remote target guardCapability Gate Internals
Runtime approvalCapability Gate Internals
Process splitProcess Split
Helper confinementPer-Process Confinement
Credential storeCore Flow
Connector call isolationProcess Split and Tools
Channel access controlChannels

One boundary has no section of its own: the PTH guard and CLI startup security. Those checks live in nanoinfra/security/ and in the CLI entrypoints.

When changing tools, channels, file access, WebUI workspace behavior, or network fetching, treat security as part of the functional behavior. Update docs if the user-facing boundary changes.

Extension Points​

ExtensionHow
ProviderAdd ProviderSpec in providers/registry.py, add a schema field in config/schema.py, implement a provider only if the generic backend is not enough, and follow development.md#adding-an-llm-provider
ChannelExport a ChannelPlugin descriptor, keep its runtime and optional setup surfaces in one package, and follow channel-package-guide.md
ToolImplement a tool under agent/tools/ or expose a plugin entry point
MCPAdd tools.mcpServers config
SkillAdd workspace skill files under <workspace>/skills/ or built-in skills under nanoinfra/skills/

Prefer existing registry/discovery patterns over ad hoc wiring.

Testing and Verification​

Common checks:

pytest tests/test_openai_api.py::test_function -v
ruff check nanoinfra/
cd webui && bun run test
cd webui && bun run build

Choose tests based on the changed surface:

ChangeMinimum useful verification
Provider behaviorProvider unit tests or a mocked API path nanoinfra agent -m "Hello!" with safe config when possible
Channel behaviorChannel tests plus nanoinfra gateway startup path
WebUI behaviorWebUI tests/build and, for routing/settings/chat changes, browser-level verification through the gateway
Tool behaviorTool unit tests and an agent-run path when schema or model-facing behavior changes
DocsLink checks, command accuracy against CLI/schema, and git diff --check

For user-facing flows, prefer at least one verification path through the public surface the user actually touches. That surface is a CLI command, an HTTP endpoint, WebSocket/WebUI, a chat channel, or a packaged import.