WebSocket Server Channel
nanoinfra can act as a WebSocket server, allowing external clients (web apps, CLIs, scripts) to interact with the agent in real time via persistent connections.
Features
- Bidirectional real-time communication over WebSocket.
- Streaming support — receive agent responses token by token.
- Token-based authentication (static tokens and short-lived issued tokens).
- Multi-chat multiplexing — one connection can run many concurrent
chat_ids. - TLS and SSL support (WSS) with enforced TLSv1.2 minimum.
- Client allow-list via
allowFrom. - Auto-cleanup of dead connections.
Quick Start
1. Configure
The WebSocket channel is enabled by default. Add only the fields you want to
override under channels.websocket:
{
"channels": {
"websocket": {
"host": "127.0.0.1",
"port": 8765,
"path": "/",
"tokenIssueSecret": "your-webui-password",
"websocketRequiresToken": true,
"allowFrom": ["*"],
"streaming": true
}
}
}
2. Start nanoinfra
nanoinfra gateway
You should see:
WebSocket server listening on ws://127.0.0.1:8765/
3. Connect a client
# Using websocat
websocat ws://127.0.0.1:8765/?client_id=alice
# Using Python
import asyncio, json, websockets
async def main():
async with websockets.connect("ws://127.0.0.1:8765/?client_id=alice") as ws:
ready = json.loads(await ws.recv())
print(ready) # {"event": "ready", "chat_id": "...", "client_id": "alice"}
await ws.send(json.dumps({"content": "Hello nanoinfra!"}))
reply = json.loads(await ws.recv())
print(reply["text"])
asyncio.run(main())
Connection URL
ws://{host}:{port}{path}?client_id={id}&token={token}
| Parameter | Required | Description |
|---|---|---|
client_id | No | Identifier for allowFrom authorization. Auto-generated as anon-xxxxxxxxxxxx if omitted. Truncated to 128 chars. |
token | Conditional | Authentication token. Required when websocketRequiresToken is true or token (static secret) is configured, unless the request comes through an authenticated trustedProxyAuth peer. |
Wire Protocol
All frames are JSON text. Each message has an event field.
Server → Client
ready — sent immediately after connection is established:
{
"event": "ready",
"chat_id": "uuid-v4",
"client_id": "alice"
}
message — full agent response:
{
"event": "message",
"chat_id": "uuid-v4",
"text": "Hello! How can I help?",
"media": ["/tmp/image.png"],
"reply_to": "msg-id"
}
media and reply_to are only present when applicable.
delta — streaming text chunk (only when streaming: true):
{
"event": "delta",
"chat_id": "uuid-v4",
"text": "Hello",
"stream_id": "s1"
}
stream_end — signals the end of a streaming segment:
{
"event": "stream_end",
"chat_id": "uuid-v4",
"stream_id": "s1"
}
reasoning_delta — incremental model reasoning / thinking chunk for the active assistant turn. Mirrors delta but targets the reasoning bubble above the answer rather than the answer body:
{
"event": "reasoning_delta",
"chat_id": "uuid-v4",
"text": "Let me decompose ",
"stream_id": "r1"
}
reasoning_end — close marker for the active reasoning stream. WebUI uses this to lock the in-place bubble and switch from the shimmer header to a static collapsed state:
{
"event": "reasoning_end",
"chat_id": "uuid-v4",
"stream_id": "r1"
}
Reasoning frames only flow when the channel's showReasoning is true (default) and the model returns reasoning content. That content comes from DeepSeek-R1, Kimi, MiMo or OpenAI reasoning models, from Anthropic extended thinking, or from inline <think> or <thought> tags. Models without reasoning produce zero reasoning_delta frames.
runtime_model_updated — broadcast when the gateway default runtime changes or
when a config reload requires clients to refresh their model catalog:
{
"event": "runtime_model_updated",
"model_name": "openai/gpt-4.1-mini",
"model_preset": "fast"
}
model_preset is omitted when no named preset is active. WebUI clients use this event
to refresh model settings after default-runtime and config changes. /model <preset>
is session-scoped. Its selection is reflected through session_updated and the
session row's model_preset field instead of this global event.
attached — confirmation for new_chat / attach inbound envelopes (see Multi-chat multiplexing):
{"event": "attached", "chat_id": "uuid-v4"}
error — soft error for malformed inbound envelopes. The connection stays open:
{"event": "error", "detail": "invalid chat_id"}
Client → Server
Legacy (default chat): send a plain string, or a JSON object with a recognized text field:
"Hello nanoinfra!"
{"content": "Hello nanoinfra!"}
Recognized fields: content, text, message (checked in that order). Invalid JSON is treated as plain text. These frames route to the connection's default chat_id (the one announced in ready).
Typed envelopes (multi-chat): any JSON object with a string type field is a typed envelope:
type | Fields | Effect |
|---|---|---|
new_chat | — | Server mints a new chat_id, subscribes this connection, replies with attached. |
attach | chat_id | Subscribe to an existing chat_id (e.g. after a page reload). Replies with attached. |
message | chat_id, content | Send content on chat_id. First use auto-attaches. No explicit attach needed. |
See Multi-chat multiplexing for the full flow.
Configuration Reference
All fields go under channels.websocket in config.json.
Connection
| Field | Type | Default | Description |
|---|---|---|---|
enabled | bool | true | Enable the WebSocket server. Set to false only when you intentionally do not want the bundled WebUI/WebSocket surface. |
host | string | "127.0.0.1" | Bind address. Use "0.0.0.0" to accept external connections. |
port | int | 8765 | Listen port. |
path | string | "/" | WebSocket upgrade path. Trailing slashes are normalized (root / is preserved). |
publicWsUrl | string | "" | Exact public ws:// or wss:// endpoint returned by /webui/bootstrap. Set this when a reverse proxy forwards requests with an origin Host header (for example, wss://claw.example.com/). Its path must match path. |
maxMessageBytes | int | 37748736 | Maximum inbound message size in bytes (1 KB – 40 MB). Default (36 MB) is sized to accept up to 4 base64-encoded image attachments at 8 MB each. Lower it if the channel only carries text. |
Authentication
| Field | Type | Default | Description |
|---|---|---|---|
token | string | "" | Static shared secret. When set, clients must provide ?token=<value> matching this secret (timing-safe comparison). Issued tokens are also accepted as a fallback. A trusted proxy assertion bypasses this requirement. |
websocketRequiresToken | bool | true | When true and no static token is configured, clients must still present a valid issued token, unless trustedProxyAuth authenticates the direct proxy peer. Set to false to allow unauthenticated connections (only safe for local/trusted networks). |
tokenIssuePath | string | "" | HTTP path for issuing short-lived tokens. Must differ from path. See Token Issuance. |
tokenIssueSecret | string | "" | Secret required to obtain tokens via the issue endpoint. If empty, any client can obtain WebSocket connection tokens from tokenIssuePath (logged as a warning). /webui/bootstrap issues tokens for local/secret-authenticated requests. Trusted-proxy requests intentionally receive no bootstrap or API token. |
trustedProxyAuth | object or null | null | No-token authorization for a directly connected upstream proxy, and the only way an approver can be a person rather than a shared token. A CIDR alone never authorizes anything: the peer must match and the assertion must be present and, on the jwt path, verified and admitted. See Identity and SSO. |
trustedProxyAuth.trustedPeerCidrs | list of CIDR strings | — | Direct TCP peer networks that may present the assertion. The check reads the direct TCP peer address only, never a forwarded-client header. IPv4, IPv6, and IPv4-mapped IPv6 peers are supported. Universal CIDRs (0.0.0.0/0, ::/0) are rejected. |
trustedProxyAuth.assertionHeader | string | — | Header injected by the identity-aware proxy after successful authentication. Routing/client metadata headers (Host, Forwarded, X-Forwarded-*, X-Real-IP, CF-Connecting-IP) are rejected, because a client influences them and none of them can name a person. |
trustedProxyAuth.assertionFormat | "jwt" or "plain" | "jwt" | jwt verifies the signature of the assertion. plain trusts a bare string on the peer CIDR alone, because a string carries no signature that anything could check. A block from an earlier version has two fields and now fails at startup. Write "plain" to keep the older behaviour and read the warning it prints. |
trustedProxyAuth.issuer | string | "" | Required on the jwt path. Must equal the iss claim exactly. |
trustedProxyAuth.audience | string | "" | Required on the jwt path. Must equal the aud claim exactly. It is what stops a token minted for another application of the same provider from authenticating here. |
trustedProxyAuth.jwksUrl | string | "" | Where the signing keys come from. Cached, and refreshed once when a token names an unknown kid, so a key rotation recovers with no restart. The address is checked by the narrow guard, which allows RFC1918 because a private provider lives there, and blocks loopback, link-local and the cloud metadata address. |
trustedProxyAuth.jwks | object or null | null | The keys inline, for a deployment that will not let the gateway make that request. Exactly one of jwksUrl or jwks must be set: two sources are two answers for one kid, and a reader of the config cannot tell which is in force. |
trustedProxyAuth.identityClaim | string | "email" | Which claim names the person. The resolved actor is webui:<claim value>, and that whole string is what gates.approvers matches. sub is available for a stable opaque id. |
trustedProxyAuth.allowedIdentities | list of strings | — | Who may reach the agent. Compared exactly against the claim value. |
trustedProxyAuth.workspaceKeyClaim | string | sub | What per-identity storage is filed under, which is not what names the person. The directory for one identity is derived from the issuer and this claim, so a person who changes address keeps their files and a reassigned address inherits nobody's. A token that verifies without this claim is admitted and shares the default workspace. |
trustedProxyAuth.signOutPath | string | — | The proxy's sign-out route on this origin, such as /oauth2/sign_out or /cdn-cgi/access/logout. The WebUI offers Sign out only when this is set, because the session is the proxy's cookie: the gateway cannot end it and will not show a button that does nothing. A path and not a URL — the schema refuses one, since a URL here would send every reader of the page somewhere else. |
trustedProxyAuth.requiredClaims | object | — | Other claims that must match exactly, which covers a whole domain through hd or a group through a mapped claim without listing every person. |
trustedProxyAuth.allowAnyVerifiedIdentity | bool | false | Admits every identity the signature verifies. A jwt block must name allowedIdentities or requiredClaims, or set this. A block that names nobody refuses to load. The gateway names this posture at every start. |
tokenTtlS | int | 300 | Time-to-live for issued tokens in seconds (30 – 86,400). |
Token Scopes
An API token carries scopes, and a route requires one. Before this, any API token was authority over every route the gateway serves. That was survivable while the WebUI was the only holder, and not survivable once a second, narrower one exists.
| Scope | What it covers |
|---|---|
operate | Everything. Implies each of the others. |
read | Reading sessions, settings and listings. |
chat | Starting a turn. |
approve | Answering a suspended action in the approvals inbox. |
secrets | The /api/webui/secrets* routes. |
A token issued without scopes holds operate, which is what the WebUI's own bootstrap token
gets, so nothing changes for a browser session. The direction that matters is the default on the
other side. A route that does not state a scope requires operate. So a route added later
without anyone thinking about scopes is operator-only rather than reachable by whatever narrow
token happens to exist.
Two surfaces state theirs today, and they are the two a narrow client must not reach. The
approvals routes require approve, because answering authorises a remote command. The secret
routes require secrets. The secret routes are scoped as a group rather than per handler, so one
added later cannot forget.
The two commissioning routes state no scope, so they require operate, and that
is the answer rather than an oversight. POST /api/webui/automations/{id}/commission
runs a model turn, and
POST /api/webui/automations/{id}/grant writes a standing grant — a permission
that covers a command on those hosts in every future unattended turn.
Answering one suspended action is approve. Authorising all of them forever is
strictly more, so it sits with operate and not beside it.
A trusted-proxy identity is deliberately unscoped. The proxy is the authority for those requests, and narrowing them here would be a second, disagreeing opinion about the same request.
Access Control
| Field | Type | Default | Description |
|---|---|---|---|
allowFrom | list of string | ["*"] | Allowed client_id values. "*" allows all [] denies all. |
Streaming
| Field | Type | Default | Description |
|---|---|---|---|
streaming | bool | true | Enable streaming mode. The agent sends delta + stream_end frames instead of a single message. |
Keep-alive
| Field | Type | Default | Description |
|---|---|---|---|
pingIntervalS | float | 20.0 | WebSocket ping interval in seconds (5 – 300). |
pingTimeoutS | float | 20.0 | Time to wait for a pong before closing the connection (5 – 300). |
TLS/SSL
| Field | Type | Default | Description |
|---|---|---|---|
sslCertfile | string | "" | Path to the TLS certificate file (PEM). Both sslCertfile and sslKeyfile must be set to enable WSS. |
sslKeyfile | string | "" | Path to the TLS private key file (PEM). Minimum TLS version is enforced as TLSv1.2. |
Token Issuance
For production deployments where websocketRequiresToken: true, use short-lived tokens instead of embedding static secrets in clients.
How it works
- Client sends
GET {tokenIssuePath}withAuthorization: Bearer {tokenIssueSecret}(orX-Nanoinfra-Authheader). - Server responds with a one-time-use token:
{"token": "nwt_aBcDeFg...", "expires_in": 300}
- Client opens WebSocket with
?token=nwt_aBcDeFg...&client_id=.... - The token is consumed (single use) and cannot be reused.
The embedded WebUI's /webui/bootstrap route returns a WebSocket token and
REST api_token for local or secret-authenticated requests. When
trustedProxyAuth authenticates the direct proxy peer, it returns connection
metadata only: no bootstrap token and no REST API token. On that path the
WebSocket handshake and the later REST requests need no token query parameter.
Trusted proxy no-token bootstrap
trustedProxyAuth lets an identity-aware reverse proxy authenticate the user
instead of a token. The fields are in the table above.
Identity and SSO holds the rest: the model, the two
requirements the proxy must meet, and a worked Cloudflare Access setup.
Do not enable this option if untrusted clients can connect directly to the nanoinfra listener.
Example setup
{
"channels": {
"websocket": {
"port": 8765,
"path": "/ws",
"tokenIssuePath": "/auth/token",
"tokenIssueSecret": "your-secret-here",
"tokenTtlS": 300,
"websocketRequiresToken": true,
"allowFrom": ["*"],
"streaming": true
}
}
}
Client flow:
# 1. Obtain a token
curl -H "Authorization: Bearer your-secret-here" http://127.0.0.1:8765/auth/token
# 2. Connect using the token
websocat "ws://127.0.0.1:8765/ws?client_id=alice&token=nwt_aBcDeFg..."
Limits
- Issued tokens are single-use — each token can only complete one handshake.
- Outstanding tokens are capped at 10,000. Requests beyond this return HTTP 429.
- Expired tokens are purged lazily on each issue or validation request.
Multi-chat multiplexing
A single WebSocket can carry many concurrent chats. The server tracks chat_id -> {connections} as a fan-out set, so the same chat can also be mirrored across multiple connections (e.g. two browser tabs).
Typical flow (web UI with a sidebar)
client server
| --- connect --------------------> |
| <-- {"event":"ready", |
| "chat_id":"d3..."} (default)|
| |
| --- {"type":"new_chat"} ---------> |
| <-- {"event":"attached", |
| "chat_id":"a1..."} |
| |
| --- {"type":"message", |
| "chat_id":"a1...", |
| "content":"hi"} ------------> |
| <-- {"event":"delta", ...} |
| <-- {"event":"stream_end", ...} |
| |
| --- {"type":"attach", | # after page reload
| "chat_id":"a1..."} ---------> |
| <-- {"event":"attached", ...} |
Rules
- Every outbound event carries
chat_id. Clients must dispatch by that field. chat_idformat:^[A-Za-z0-9_:-]{1,64}$. Non-matching values returnerror.messageauto-attaches on first use — no separateattachis required for chats the server minted (new_chat) on the same connection.- Errors (invalid envelope, unknown
type, badchat_id) are soft: the server replies with{"event":"error","detail":"..."}and keeps the connection open.
Backward compatibility
Legacy clients that only send plain text or {"content": ...} keep working unchanged: those frames route to the connection's default chat_id (the one from ready). No config flag is needed.
Security boundary
chat_id is a capability: anyone holding a valid WebSocket auth credential and the chat_id can attach to that conversation and see its output. This is safe for nanoinfra's local, single-user model. Multi-tenant deployments should namespace chat_ids per user (or introduce a per-tenant auth gate) — nanoinfra does not do this today. Multi-Tenancy is the page that covers what to do instead.
Security Notes
- Timing-safe comparison: Static token validation uses
hmac.compare_digestto prevent timing attacks. - Defense in depth:
allowFromis checked at both the HTTP handshake level and the message level. - chat_id as capability: see Multi-chat multiplexing. Auth on the WebSocket handshake is the single line of defense. Callers who pass it can attach to any chat_id they know.
- TLS enforcement: When SSL is enabled, TLSv1.2 is the minimum allowed version.
- Default-secure:
websocketRequiresTokendefaults totrue. Explicitly set it tofalseonly on trusted networks.
Media Files
Outbound message events may include a media field containing local filesystem paths. Remote clients cannot access these files directly — they need either:
- A shared filesystem mount, or
- An HTTP file server serving the nanoinfra media directory
Common Patterns
Trusted local network (no auth)
{
"channels": {
"websocket": {
"host": "0.0.0.0",
"port": 8765,
"websocketRequiresToken": false,
"allowFrom": ["*"],
"streaming": true
}
}
}
Static token (simple auth)
{
"channels": {
"websocket": {
"token": "my-shared-secret",
"allowFrom": ["alice", "bob"]
}
}
}
Clients connect with ?token=my-shared-secret&client_id=alice.
Public endpoint with issued tokens
{
"channels": {
"websocket": {
"host": "0.0.0.0",
"port": 8765,
"path": "/ws",
"tokenIssuePath": "/auth/token",
"tokenIssueSecret": "production-secret",
"websocketRequiresToken": true,
"sslCertfile": "/etc/ssl/certs/server.pem",
"sslKeyfile": "/etc/ssl/private/server-key.pem",
"allowFrom": ["*"]
}
}
}
Custom path
{
"channels": {
"websocket": {
"path": "/chat/ws",
"allowFrom": ["*"]
}
}
}
Clients connect to ws://127.0.0.1:8765/chat/ws?client_id=.... Trailing slashes are normalized, so /chat/ws/ works the same.