Changelog
Releases before 1.0.0 are recorded in long form on the Release Archive, which is the only account of the 0.x line.
All notable changes to nanoinfra are recorded here.
The format follows Keep a Changelog 1.1.0, and this project adheres to Semantic Versioning.
A line names an effect a user can observe and carries a reference. The reasoning behind a change — what was measured, what was wrong first — lives in the issue or the proposal that line points at, not here.
Unreleased
2.3.0 — 2026-09-08
Added
- The composer says how much of the context window a thread is using, and what the last eight provider calls cost. Each bar splits the call's input into the three buckets billed separately, so a warm prompt and a cold one of the same size no longer look identical.
Fixed
-
A streamed segment carries the cost of the provider call behind it. The value reached the channel's own tests but never a real turn: the loop wraps the delivery callback, and the hook offers a fact only to a callback that names it — so 51
stream_endrecords across 13 sessions arrived with no usage at all. -
A
socks://proxy configured for the OpenAI-compatible or xAI provider now connects. httpx knows no such scheme, so on a host carrying that spelling every model call failed while its HTTP client was still being built. -
A tool-call id issued by an earlier response is no longer replayed as a Responses input item id. After a restart, a model switch or any provider-state mismatch the endpoint could reject the whole request as an item that does not belong to this connection.
-
A Grok hosted search the stream abandoned no longer reads as a finished answer. The request is retried instead of answering a search question from nothing, and the search's activity row stops saying "searching" forever.
-
A relative
working_dirresolves against the workspace.execran the command in whatever directory the gateway process was started in, and a CLI app refused the same path as "outside the configured workspace" — one input, two wrong answers in opposite directions. -
Two MCP tools with non-ASCII names both stay registered and callable. Sanitization erased everything that told
mcp_weather_获取天气apart from its neighbour, so the second tool replaced the first while the connect log still counted both. -
AGENTS.md, SOUL.md and USER.md are bounded in the system prompt, and say so when they are shown in part. They were embedded whole on every turn, and Dream is the writer of two of the three.
-
The idle-session summary cache is bounded. Entries were dropped only when a session was reopened, so a session archived and then abandoned held its slot for the life of the process.
-
A Codex chat now keeps one prompt cache key for its whole length, so the prefix it already paid for is billed as a cache read instead of being re-tokenised every turn. (#203)
-
Codex builds its TLS context once per process rather than once per request, so streaming no longer stutters while the CA bundle is read again before each call.
-
A background task that raises is now logged with its traceback and the name of the coroutine that failed. Memory consolidation and idle auto-compaction run through that path, so either could have been failing on every turn with nothing but asyncio's unattributed "Task exception was never retrieved" to show for it.
-
Fourteen log calls that passed
exc_info=Truenow record the traceback. loguru takes keywords as format bindings, so the flag was accepted, unused and dropped — every one of those sites believed it was keeping an exception's traceback and was keeping none. -
Six log messages that used printf
%snow interpolate. The message printed a literal%sand the argument — an entry-point name, a failing tool's name — was dropped. -
A provider backoff is visible on any channel that already takes progress updates. The event was published and then discarded before dispatch, so outside the CLI a two-minute
Retry-Afterlooked like a chat that had simply gone quiet. -
A long Anthropic answer is no longer cut off at the stream idle timeout. The 90 s bound measured total generation instead of silence whenever no streaming callback was attached — which is every retry after a stall. (
4de728a5) -
Guidance queued in one chat is no longer sent into another. Opening a second chat that was idle read as the first chat finishing, so the waiting prompt went to whatever conversation was on screen; it now stays in its own chat until that chat's run ends.
-
One slow WebSocket client no longer delays the frames owed to every other client on the same chat, and a client that stops reading is disconnected rather than allowed to grow an unbounded outbound backlog in server memory.
-
An email whose hand-off to the agent fails stays unread and is delivered again on the next poll. It was marked
\Seenand deduped during the fetch, so a failed hand-off was never retried and the mailbox reported it as handled. -
A filtered email — self-sent, failing SPF/DKIM, or not allow-listed — is left unread. Marking it read made the mailbox's unread state stop reporting what the bot processed;
postActionstill decides what happens to it. -
A degraded consolidation no longer loses the messages past the first 16,000 characters of its raw dump. When the summarising model errors or runs out of room, the batch is dumped to history as
(part i/n)entries instead of one truncated entry — the compaction cursor advances past that batch either way, so the part that used to be cut off left the session with no copy anywhere. (#109) -
Editing a cron job no longer cancels the job that is running. The scheduler's timer task is the task running the turn, and every edit re-armed that timer — so the agent's own cron tool, an operator toggling an automation in the WebUI, and a commissioning verdict each killed the turn mid-run and left the job to fire again. (upstream PR 5686)
-
**in afind_filesorgrepglob now spans any number of directories, including none. The example both tools hand the model,tests/**/test_*.py, matched only paths with exactly one directory in between, andsrc/**/*.pymatched nothing at all. (upstream PR 5692) -
A consumer slower than the agent no longer loses an event from an SDK stream. Closing the stream evicted the oldest queued event to make room for its end marker, which in practice cost the final
text_completed. (upstream PR 5635) -
A backgrounded nanoinfra process now writes its
print()output to its log. Redirected output is block-buffered, so a child that hung or was killed left an empty log file, exactly when it was being read to find out why. (upstream PR 5412) -
The retry that drops an image and succeeds is counted. Only the failed attempt was recorded, so the tokens the answer actually cost were charged to nobody. (#176)
-
A model that raises instead of answering now fails over. An unauthenticated GitHub Copilot, or an endpoint that refuses the connection, skipped every configured fallback and was not retried.
-
A dropped or reset connection is retried.
ConnectError,ReadErrorandRemoteProtocolErrorwere reported as errors of no known kind, which made a transport failure look permanent. -
An OpenAI
server_erroris retried instead of ending the turn on the first attempt. -
A primary with an open circuit and no usable fallback returns the primary's own error,
Retry-Afterincluded, and asks to be retried once the cooldown is over.
Changed
nanoinfra webuistarts even when model setup is incomplete, on every run and under--yes, naming the provider at fault. The second run used to refuse, which shut a half-configured install out of the Settings → Models screen that repairs it.
Removed
-
The CLI Quick Start offer from
nanoinfra webui. Provider and model setup is finished in WebUI Settings → Models. -
An SDK run started with
ephemeral=Truewrites nothing. The turn was persisted like any other — session file, mid-turn checkpoints and the cached session object — while only thesession_turn_persistedevent was withheld.
2.2.8 — 2026-09-07
Changed
mcpis capped below 1.30. That release changes three defaults at once — redirects are followed only within the endpoint's origin, idle Streamable HTTP sessions expire, and the OAuth client validates the authorization server'sissuer— and the first two govern how this project talks to an MCP server over HTTP.
Fixed
- A command in the
bwrapsandbox now runs nanoinfra's own Python. It resolved to the base interpreter instead, so a dependency this project declares and installs —openpyxl, and anything else a shell command imports — failed inside the sandbox while sitting installed. (#276) - A redirect to
/dev/nullno longer counts as a workspace-bypass attempt. Three of them in one turn — which any non-trivial shell command reaches — told the agent it had hit a hard policy boundary about a path it did not want, and the turn derailed. (#282)
2.2.7 — 2026-09-06
Added
- Metrics → Calls filters by source and actor. Both were already on every row and in the payload with no control to filter on them, so finding what one approver authorised, or what one channel ran, meant reading the table by eye. (#274)
2.2.6 — 2026-09-06
Changed
Abilitiesin the sidebar starts closed, and opens only when you click it. It opened by default, and it also reopened itself whenever the active page was one of its two members — so a visit to Skills undid the collapse. Both are gone: absent means closed, and the heading is the only thing that opens it.Infrastructurealready started closed and keeps its own reopen behaviour, which was not part of the ask. (#253)- Both rail groups remember whether you closed them. They held that in local state, so the choice
lasted until the next reload — and the comment on that state called collapsing "the operator's
choice, not the default", which a choice that does not survive a reload is not. They now use
collapsed_groups, the map the sidebar already round-trips for chat project groups, undernav:-prefixed keys so they cannot collide with a project of the same name. (#253)
2.2.5 — 2026-09-06
Fixed
- A row's detail in Metrics → Calls opens under that row instead of after the whole table. Opening row 1 of 100 put its fields below row 100, so reading a call meant scrolling the page away from the row that was clicked. The chevron on the left already promised an inline disclosure, and a disclosure that opens a hundred rows away is a broken affordance rather than a layout preference. Thirteen fields also lay out in three columns on a wide screen now, so the rows below do not travel far. (#274)
2.2.4 — 2026-09-06
Added
- The Live tab draws its numbers as well as printing them. Each gauge grows a sparkline of the last three minutes, kept in the tab — nothing stores gauge history, because the gauges are sampled when read, and the panel says so rather than implying it holds yesterday. A gauge that could not be read gets its dash and no plot: an empty plot area reads as flat at zero. (#274)
- Context used and its limit are one meter instead of two tiles. They were never two facts, and reading 428K against 1,048,576 is arithmetic the panel should do. The fill carries severity past 75% and 90%, with the percentage always spelled out — a status colour never carries meaning alone. An absent or zero limit reads as unknown rather than as 0%. (#274)
- Two charts under the gauges: calls per minute, derived from the difference between
successive reads of the cumulative counters the way
rate()does, and a latency histogram over the nine buckets. An empty latency bucket keeps its row, because a band with no calls is information.GET /api/webui/metrics/countersserves both. (#274)
Fixed
- The counters charts cannot take the Live tab down. They render inside it, so trusting the route's shape meant an older gateway — or any unexpected answer — unmounted every gauge above them. The same failure the scale row had one release earlier. (#274)
2.2.3 — 2026-09-05
Added
- An
Approvalstab in Metrics: how many actions the gate held for a person, how many they answered, how many they refused, how many expired, and the median time to answer — the number that says whether the gate is working rather than merely running. A person's refusal is counted apart from a policy refusal nobody was asked about, because merging them overstated the approver's denials sixteen-fold on the deployment this was measured against. An ask that was neither answered nor expired is named on its own: nothing ran and nothing said why. (#274) - A scale row at the top of Metrics → Usage: servers, skills, agents, MCP servers and connectors,
in one request. Every one of these existed and was scattered across five settings pages. A
count that cannot be read shows
—and names itself rather than reading as zero. (#274) /metricsexports counters and a latency histogram, so a Prometheus install can ask for a rate and a quantile rather than only a level:nanoinfra_llm_calls_total,nanoinfra_llm_tokens_total(by input, output, cache read and cache write),nanoinfra_tool_calls_total, andnanoinfra_llm_duration_msover nine buckets. Accumulated in memory for the life of the process, because a count over a table with a purge is not monotonic and arate()over a falling counter is nonsense. (#274)- Two process-health gauges,
nanoinfra_rss_bytesandnanoinfra_event_loop_lag_ms. A rising lag is a blocked loop, and neither number has an event to be driven by, so both are sampled at read. (#274)
Fixed
- The gate audit viewer no longer offers
denyas a decision filter. Nothing writes it: the name lives on as the outcome enum and as the operator socket's wire verb, and the log records that outcome asdeniedso it speaks the operator's vocabulary. The filter could only ever match zero records. (#274)
2.2.2 — 2026-09-05
Fixed
- The socket directories are setgid for real this time, verified in the built image rather than
reasoned about. v2.2.1 set the mode before the
chowninprepare, which was right and not enough: the post-bind block re-applies the mode, and by then the directory already carries the helper's group, so itschmod 2710dropped the bit again and returned success. All five directories were still at710on 2.2.1. Oneset_socket_dir_modehelper now owns that mode and takes the directory back to root before setting it, which is the only order that works on a fresh directory, on a re-apply, and on the0700one the Python side creates.
2.2.1 — 2026-09-05
Fixed
- Every helper socket directory is actually setgid now, so a socket keeps its shared group when
the helper rebinds it.
chmod 2710ran afterchown, and withoutCAP_FSETID— which the published compose file does not grant — that silently drops the setgid bit and returns success, so all five directories sat at710while the code's own comments described2710. The operator socket is the one that paid: the executor deliberately does not joinnanoinfra-op, so the inherited group is the only mechanism it has, and its absence left the racy root chown as the only thing setting it — plus a[Errno 1] Operation not permittedon every boot. apply_socket_groupsets the mode even when it cannot set the group. The two were in onetry, so a refused chown skipped the chmod — and the mode is the half that grants a peer its write bit.
2.2.0 — 2026-09-05
Added
- A
Metricsdestination in the rail, with three tabs.Usageshows spend per model over a chosen window of 7, 30, 90 or 365 days, with cache writes, truncated answers, time to first token, wall clock and a per-model cost — five of which the store has recorded since 2.0.0 and none of which reached a pixel.Liveshows seven point-in-time gauges, starting with how many suspended actions are waiting for a person.Callsreads thetool_callstable, which had a writer, a pruner and a purge log and no reader at all. (#235, #232) pricingin config gives a model four rates in USD per million tokens — input, output, cache read and cache write — keyed"<provider>/<model>". Until one is set the Usage tab says no prices are configured rather than showing a spend of$0.00. (#235)gateway.metricsEnabledserves the same gauges in Prometheus text format at/metricson the gateway's own port. Off by default.gateway.metricsTokensets the bearer token a scrape must present; with no token only a loopback bind is served, so enabling metrics on a port a reverse proxy fronts does not publish them. (#235)- The Usage tab lists why calls failed, per error kind, status code and provider. The failure count was already shown; the reason behind it was recorded and never read. (#235)
- The Usage tab breaks the window down by what started the turns — chat, API, automations, memory,
system.
sourcewas aggregated per day and readable only inside one heatmap cell's tooltip, so "what does automation cost me this month" meant opening thirty tooltips. (#235) - A model row in the Usage tab expands to the four measurements the columns are checked against: how many calls the provider reported against how many were tokenized locally, output tokens per second of generation, measured against reported output, and the streamed split that time to first token is averaged over. (#235)
Fixed
- A cached token is no longer billed twice.
prompt_tokensis the logical input and includes the cached halves, so charging it at the input rate and the cached count at the cache rate over-stated a warm cache badly — 5.7× on published Kimi K3 rates for a 90%-cached prompt. The three input buckets are now disjoint. (#235) - A usage row names the provider that was configured rather than the class that made the call.
OpenAICompatProviderserves every OpenAI-compatible API, so Moonshot, DeepSeek, Groq, OpenRouter and forty others were all recorded asopenaicompat— which left "which provider is expensive" unanswerable and meant a rate set formoonshot/kimi-k3could never match its own rows. Single-provider backends were wrong too:openai_codexrecordedopenaicodex. (#235) - Both spellings of the pricing rate keys are accepted (
inputPerMtokandinputPerMTok), and an entry that states no rate nanoinfra recognises reads as unpriced rather than as$0.00. A mistyped key used to produce a confident zero over a month of real spend. (#235)
Changed
- Model rates are set where the model is:
Settings → Models → Pricing, with a live figure showing what the recorded window would have cost at the rates typed, a warning when the model reads cached tokens and the cache rate is still zero, and a note when two configurations name the same model and therefore share one bill.Settings → Providers → Default pricingsets rates for a whole provider, including a free checkbox — the one-edit answer for a local fleet. (#235) - The Settings overview no longer opens with the token summary and the heatmap. They are the
Usage tab now, and the overview keeps one row that leads there — which also ends the
five-second re-aggregation of
llm_callsthat a settings page left open used to run. (#235)
2.1.0 — 2026-09-04
Added
- A tool group, MCP server or connector can be set to
attach: "search": its schemas leave every prompt and the model loads them itself by calling the newtool_searchtool, with one shared pointer in place of a per-item advertised line.mentionstays the user-driven deferral; both respect the acting agent's ceiling, which neither can widen past.
Fixed
- A tag no longer publishes to PyPI or GHCR before the Test Suite has finished on that commit. v2.0.0 shipped a wheel that could not be imported on the minimum supported Python because the publish jobs are quicker than the tests and nothing connected them. (#270)
Changed
- CI builds the image with Buildx and the Actions cache, the way the publish workflow already did. (#271)
2.0.4 — 2026-09-04
Fixed
@agent:<name>in a message answers as that agent. The composer offered the token, completed it, and then ignored it: the turn ran as the deployment default, so naming an agent looked like it did nothing. (#269)
2.0.3 — 2026-09-04
Fixed
- The WebUI inbox is no longer warned about as an unreachable chat channel. A deployment whose approver sits there was told on every poll that a suspended action reaches nobody, while the inbox had the request. (#267)
2.0.2 — 2026-09-04
Fixed
- The
Agentsdestination is in the sidebar whether or not the deployment names an agent. It was gated on the roster, so on a fresh install nothing led to the page that configures the agent answering every turn. (#266)
Changed
Abilitiesis no longer gated on the roster either, and it opens expanded — one sidebar shape whatever the deployment holds, with Apps and Skills still one click away. (#253)
2.0.1 — 2026-09-04
Fixed
- nanoinfra imports on Python 3.11 again. A dataclass field on
RequestContextdefaulted to amappingproxy, which 3.11 refuses for any default whose type is unhashable — so 2.0.0 failed at import on the minimum supported version. (#266)
2.0.0 — 2026-09-04
Added
- One approval can become a standing grant. The grant is derived from the payload the executor actually rendered, defaults to expiring, and asks once more before it never does. (#217)
- Every server keeps notes the agent and the operator both write, and a turn that names a server reads them. A note does not expire; one that disagrees with what you see is evidence the infrastructure changed. (#222)
- A queryable record of every tool call: which tool, by whom, what the gate decided, and how it ended. It stores the address of the arguments in the session history, never the arguments. (#231)
- The agent can search your own documents and answer with citations. Drop files in
<workspace>/knowledge/; nothing reaches a prompt on a turn that does not ask. (#237) - A deployment can name more than one agent, each with its own model, tools, skills and
instructions. An empty
agents.namedis exactly the single agent every deployment has today. (#247) - An agent can hand one task to a peer and wait for its answer. Membership in
agents.named[x].delegatesis the grant, and delegation is one level deep. (#250) - A delegated action records the human who asked, the agent that delegated and the peer that acted, so a reader can answer "who authorised this" without opening a second file. (#251)
- Every assistant turn says which agent answered it, beside the turn's cost. (#248)
- The composer offers the agents a message may ask for, as
@agent:<name>. The token stays in the text, because it is a preference the answering agent reads and not an invocation. (#255) - A turn that delegates shows its plan as one object in the thread — a row per peer with its outcome and its own cost — and a reload shows the same plan. (#252)
- An Agents destination lists the agents a deployment names, and an Abilities grouping collects Apps and Skills in the menu. (#253)
- Each agent's Prompt tab shows the prompt's sections with the permission on each, and the addendum that appends after them. (#256)
- An automation can name the agent it runs as, and that agent is the ceiling: its tool groups cap the turn and its skills bound the job's own picker. (#257)
- The approvals inbox names the agent that will act, and the agent that delegated to it. (#258)
- Agents are created, edited and deleted in the browser. Each one gets a page with tabs — model, tools, skills, delegates, prompt — and every binding is picked from what this deployment has. (#262)
- The deployment's own agent is one more agent: same page, same tabs, editable down to its skills, MCP servers and delegates. It still cannot be deleted. (#265, #266)
- The Prompt tab edits the prompt. The three sections that are prose can be replaced, each with the text in force shown and what replacing it costs said before you do. (#256)
- Settings → Prompts reads and writes the two prompts that run unattended,
dreamandevaluator. (#264) - An agent's tool groups, skills, MCP servers and connectors now narrow every turn it answers, not only a scheduled one — so choosing an agent is how a conversation stops paying for every server and skill installed. (#266)
Changed
- Where a deployment names agents, the composer chooses an agent instead of a model — the model belongs to the agent. A deployment that names none keeps its model selector unchanged. (#254)
- The prompt's safety notes are their own fixed section, so replacing the runtime section can no longer delete them. (#256)
- A replaced prompt section is still named in the prompt manifest and marked as overridden — a record that hid a replacement would make two different prompts look identical. (#256)
- The executor protocol is version 7, carrying the delegation chain. A deployment running the executor as a separate process must restart it alongside the gateway. (#251)
- Nothing ships a model.
agents.defaults.modelwasanthropic/claude-opus-4-5, so an unconfigured deployment looked configured on a provider it had no credential for. The first model configuration a deployment adds is now the primary one, and it answers until something else is chosen. (#266) - An agent's empty binding list means none of them, not all of them. Declaring nothing is what means everything, and the two are now separate answers everywhere: config, the tool filter and the picker. (#266)
- A persona can write
{{ agent_name }},{{ agent_role }}or{{ agent_description }}inSOUL.mdand each agent fills it in. Any other{{ }}survives verbatim. (#265)
Removed
- The config warning about a dead
agents.defaults.model. The turn fell back to the primary preset either way, so the line reported a difference nothing acts on — and the field ships empty now. (#205)
Fixed
- A system job kept its state across a restart.
dreamandevaluatorwere re-registered as brand new on every boot, so on a deployment that restarts they never ran. (#263) - An agent's addendum and its replaced prompt sections reach the model. Both were stored, shown, editable — and never passed to the prompt builder by its only call site. (#265)
- A tool refused for want of a secret says when the encryption key is the wrong one, instead of failing silently. (#217)
- A trigger whose agent declared no tool groups is no longer capped to the ungrouped tools. (#266)
- The default agent has a card on a fresh install. It hid until the deployment named an agent, so the agent that answers every turn was reachable only after naming one you did not want. (#266)
- A deployment with no model at all says so at boot and on the settings page, instead of failing
a turn with
No provider is configured for model ''. (#266)
Security
- A delegated turn can no longer reach a capability class the turn that spawned it did not hold, and a capped turn hands that ceiling to anything it spawns. (#251)
1.9.1 — 2026-09-03
Removed
- The sidebar no longer lists conversations from other channels, reverting 1.9.0. Sessions that
never held a conversation — an empty WhatsApp session,
cli:direct— showed up asNew topicrows and buried the chats, which is worse than the problem it set out to fix. (#216)
1.9.0 — 2026-09-03
Added
- A conversation held over the API — or from a chat channel, or by an automation — appears in the WebUI sidebar with its channel badged, and opens as a readable thread built from the session history. Read-only: the composer still refuses a session it does not own, and delete, file-preview and automations stay closed to other channels. (#216)
1.8.0 — 2026-09-03
Added
api.enabledlets the gateway serve/v1onapi.portitself: one process and one agent loop instead of two, sharing the MCP and connector hosts it already booted. Defaultfalse, so upgrading opens no port.nanoinfra serveis unchanged. (#214)- The API server logs one line per request — method, path, status, duration, and why a request was
refused — and its own logger stays audible without
--verbose. Never the body or theAuthorizationheader. (#215)
1.7.4 — 2026-09-03
Fixed
nanoinfra servedrains the agent's outbound event bus, so a turn that emits more than a thousand progress or stream events finishes instead of stalling mid-flight and leaving its HTTP request unanswered.gatewayandnanoinfra agentalready drained; the API server did not. (#211)
1.7.3 — 2026-09-02
Fixed
- Both API routes accept a client that resends its own transcript: what follows the last assistant
message is the turn, and earlier messages are dropped as the client's copy of a history the
server already keeps. A
systemmessage beside the prompt is joined into the turn rather than refused, so a Responses or Chat Completions client no longer gets a 400 on its first request. (#211)
1.7.2 — 2026-09-02
Added
- The prompt breakdown says which request it describes when a turn made more than one, and names the largest request the turn reached. (#208)
- An activity cluster names what the provider calls inside it cost — input, cache share and output — beside the duration it already showed. The cache share counts only the calls that reported one. (#208)
tools.groupsdeclares groups of built-in tools with thealways | mentionmodes an MCP server and a connector already had, so@diagramscan load 2,438 tokens of diagram schemas only for a turn that asks for them.diagramsandserversare predefined; both default toalways. (#210)POST /v1/responsesanswers the same agent over the Responses wire, so a client that defaults to that protocol no longer gets a 404. The caller'stoolsandinstructionsare ignored: the agent runs its own tools behind the capability gate. (#211)agents.defaults.midTurnMessagesdecides what happens to a message that arrives while a turn is running. The new default,queue, gives it a turn of its own with its own answer;injectkeeps the previous behaviour of folding it into the turn in flight. (#209)
Changed
- A message sent while the agent is still working now gets its own turn and its own reply, instead of being folded into the turn already running and answered only as part of it. (#209)
Fixed
- Each activity cluster reports its own duration instead of the whole turn's, so eight consecutive steps no longer all read the same figure. (#208)
1.7.1 — 2026-09-02
Added
/compactarchives a session's history on request instead of waiting for an idle timer or a budget threshold, and reports how many messages it archived and how many stay raw. (#212)
1.7.0 — 2026-09-01
Added
- OpenAI requests carry
prompt_cache_key, one key per chat, so its automatic prefix cache is reached instead of every session landing in the bucket its first 256 tokens hash to. Opt-in per provider: it is a body field, and a provider that rejects an unknown one answers 400. google_calendar_freebusyanswers whether a calendar is busy in a range — busy blocks only, no event detail — so availability and slot-finding no longer mean reading every event. A connector operation can now declareread_via_postfor a read whose query needs a POST body; the invariant that a POST is a write otherwise still holds.google_calendar_update_eventchanges an event by id. A partial PATCH, so an omitted field keeps its value — sending onlystart/endmoves an event and keeps the rest.mutate.remote.google_calendar_delete_eventremoves an event by id. It carriesmutate.remote, the same class as creating one, so it asks a person in an interactive turn and needs a standing grant to run unattended.
1.6.1 — 2026-09-01
Fixed
- Every turn of one chat reaches the same xAI server, so its prompt prefix can be cached.
x-grok-conv-idwas a fresh UUID per request, which announced each call as a new conversation.
Added
CHANGELOG.md, in Keep a Changelog form, covering every version since 1.0.0. A release body is now this file's section for that version.- A tag whose version has no changelog section fails the release workflow.
1.6.0 — 2026-09-01
Added
- The prompt breakdown's tool count opens the list of tools behind it, each with its size, largest first. (#203)
Changed
- A tool row in the breakdown reads as
google-calendarwith aconnectortag, rather than as the rawconnector:google-calendar. (#203)
1.5.2 — 2026-09-01
Fixed
- Every reference prefix keeps a row in the mention menu. The palette reserved two, so a deployment
with a data connector —
server:,diagram:,calendar:— lost one. (#204)
1.5.1 — 2026-09-01
No user-visible changes. A route test pinned the arguments the marketplace search is called with, and 1.5.0 added one.
1.5.0 — 2026-09-01
Requires skills-server v0.4.0.
Added
- Apps → Connectors browses the catalog and installs from it. Each row names every operation with its capability class, the hosts a token could reach, and the scopes it would carry, before the install button. (#207)
- An install reports what is still missing: a connector has no credential and no entry in
connectors.activeuntil an operator adds them. (#207)
Changed
- The marketplace reads the catalog's
kindand installs each package where its own subsystem reads it —skills/,plugins/,connector-packages/. The kind is asked of the catalog and never taken from the caller. (#207)
Fixed
- A connector already on disk read as not installed, because
installedwas asked of the skills loader for every kind. (#207)
Security
- Workspace containment runs before anything walks the install destination. A symlinked
skills/was read through before the boundary check. (#207)
1.4.0 — 2026-09-01
Added
attach: "always" | "mention"on a data connector. Amentionconnector stays active with its credential and grants, the prompt carries one line naming it, and its operations load only for a turn that names it —@<name>, or a@<kind>:<id>object of one of its kinds. (#204)- An automation declares
connectors: [...], since an unattended turn types no@. (#204)
Fixed
- The composer palette silently dropped candidates whose kind was missing from its group list. (#204)
- A paused MCP server named in message text was still tokenised and still sent. (#206)
1.3.2 — 2026-09-01
Fixed
- The connector mention prefix reads the cached object listing, so
@calendar:is in the menu on the first keystroke instead of waiting on a live API call that took seconds and failed silently. (#204)
1.3.1 — 2026-09-01
Fixed
- A paused MCP server is no longer offered in the
@palette. Picking one read as an attachment on a server nanoinfra never connected. (#206)
1.3.0 — 2026-09-01
Added
attachon an MCP server. Amentionserver is advertised in one line — name, tool count, how to attach — roughly 50 tokens against a couple of thousand, with its schemas sent only to a turn that names it. (#204)- An automation declares
mcpPresets: [...]. (#204) - Config that is ignored says so, at boot in the log and on the Models page: a dead
agents.defaults.modelbeside amodelPresetthat overrides it, and a selected preset whose provider has no credential. (#205) - The Apps page's tool count opens the tool list.
Changed
- The Apps page opens on Ready rather than the CLI tab.
Fixed
- A paused row counts its saved allowlist instead of reporting
0 toolsfor a server holding fifteen.
1.2.3 — 2026-09-01
Fixed
- Pausing an MCP server stops it connecting.
AgentLoop.from_configseeded the live map from the unfiltered config, so a paused server reconnected on the next restart. (#206) - A plugin-declared MCP server reached the loop only after a hot reload. (#140)
1.2.2 — 2026-09-01
Changed
- The MCP pause is a switch on the row; the row's checkmark becomes an explicit
⋯menu, and the second line names what the server costs instead of naming its transport. (#206)
1.2.1 — 2026-09-01
Added
- An MCP server can be paused without losing its configuration —
enabled: false. Its command, arguments, environment, headers andenabledToolslist stay; its schemas leave every prompt. (#206)
1.2.0 — 2026-08-31
Added
- A per-turn prompt breakdown, collapsed under each turn: where the input tokens went, by section and by tool source, largest first. Recorded while the prompt is assembled, because the attribution is gone by the time a request reaches a provider. Names and sizes only, never content. (#203)
1.1.4 — 2026-08-31
Added
- Settings leads with the token numbers: 30-day tokens, calls, failures, today, the measured share, the peak day, and a per-model breakdown with time to first token. (#176, #177)
Fixed
- A migrated day topped the per-model breakdown with a day's worth of tokens against 19 "calls". It counts in the day totals and not in a breakdown it cannot answer.
- Every fallback row was labelled
fallbackrather than the provider that actually answered.
1.1.3 — 2026-08-31
Fixed
- A connector's stale
last_errorcan be retracted. A successful test or a successful listing clears it; every success had been passing an empty value, which the merge ignores by design.
1.1.2 — 2026-08-31
Fixed
- Every connector call failed with
[Errno 13] Permission deniedon any deployment that had activated one. Marketplace package discovery pointed at the connector state directory, which the executor cannot traverse, andPath.is_file()propagates a permission error rather than answeringFalse. - The connector row in Apps spans both grid columns, so its name is readable and its capability classes read as chips rather than as a vertical wall.
Changed
- Connector packages get their own root,
<workspace>/connector-packages/. A package is read by three accounts; a state file is written by one.
1.1.1 — 2026-08-31
Added
- A connector against a public API runs with no credential —
credential.kind: "none"activates with no binding, mints nothing and sends noAuthorizationheader. It is still gated per operation. examples/connectors/hello-world/: onereadagainst a public API, no credential, no setup. Ships in the sdist, because apipuser has no checkout to copy from.
1.1.0 — 2026-08-31
Added
- Every finished turn shows its cost in the footer — tokens in, out, cache share and latency — persisted beside the latency so a reloaded thread shows what the live turn showed.
- One content-free row per provider attempt in
~/.nanoinfra/llm-usage.sqlite3, retained 400 days. No prompts, no responses, no reasoning, no tool payloads, no session keys; a test asserts the schema has no column content could go in. - A connector can arrive from the catalog as a declarative
connector.jsonwith nothing importable, and its calls run in a fourth confined process (nanoinfra-connector, outbound 443 only) whose group holds the executor and not the agent.
Changed
LLMResponse.usageand both hook contexts carryLLMUsage | Noneinstead of a dict.to_turn_dict()is the projection the OpenAI-compatible API, the SDK and the WebUI read.~/.nanoinfra/webui/token-usage.jsonis migrated at first start and renamed, not deleted.
Fixed
- Three usage-accounting defects the typed contract exposed: a cache count present on one call and absent on another summed as though both were measured; two calls carrying 40k of context each read as a turn carrying 80k; the over-budget finalisation path mutated the accumulator through its argument.
- Three call sites reached a provider without passing through the agent loop — WebUI title generation, the evaluator, Dream consolidation — so their tokens were charged by the provider and counted by nothing.
Security
- A connector credential is bound to the hosts it may address, checked at activation with both
hosts named. A package declaring Google scopes and its own
baseUrlwould otherwise receive a live Google token and a process that forwards it.
1.0.6 — 2026-08-31
Fixed
- A secret record written through the executor took the writer's primary group, so a group-readable
file in the wrong group answered
EACCESto the gateway: a create succeeded and the next listing raised. The store sets the group explicitly when the directory shares one.
1.0.5 — 2026-08-31
Added
- The Secrets page writes on a container deployment. Protocol version 6 carries a third request kind: the gateway encrypts because it holds the key, and the executor writes ciphertext it cannot read. Three verbs — create, update, delete — and no read.
- Connector consent completes in the browser, returning to this deployment's own origin with the credential and the activation written. A consent that fails writes nothing.
- Reload connectors reconciles the running registry with config, without a restart.
Fixed
- The no-tools request path had no timeout while the tool-carrying path wrapped the same call, so a stall held the per-session lock forever.
- A truncated consolidation was accepted as history, taking whatever it had not reached with it.
find_filesandgrepbounded their results and nothing about their scan, and walked inside an asyncexecute— one call on a slow filesystem held the event loop for every session. Both carry an entry budget and a wall clock, and a partial answer says so.- A cron job froze the workspace path one person's turn resolved, naming their identity directory,
into
jobs.json.
Security
- A Slack file download followed any URL the workspace handed it, redirects included, with no SSRF validation and no DNS pinning.
1.0.4 — 2026-08-30
Added
- Data connectors. A connector reaches one data source with a capability class per
operation, so reading a calendar and writing to it are two decisions — where an MCP tool
declares nothing and every one of them resolves to the fail-closed
mutate.remote. google-calendarships first, with threereadoperations and onemutate.remote.- The call runs in the executor: protocol version 5 carries a second request kind, and the method, path, class and scopes come from the installed manifest — so a frame cannot describe a call the package never declared, and the agent process holds no token.
- A standing grant can name a connector, so an unattended write is expressible rather than permanently denied.
1.0.3 — 2026-08-29
Fixed
- The test that pinned the previous dependency policy now states the current one. v1.0.2 moved four packages into the base install and left an assertion saying the opposite, so the release shipped and CI failed.
1.0.2 — 2026-08-29
Changed
asyncssh,ansible-runner,boto3andaiohttpare base dependencies. The container image installs no extras, so the published image came up as an agent for infrastructure that could not reach infrastructure.nanoinfra[servers]andnanoinfra[api]still resolve.
1.0.1 — 2026-08-29
No behaviour changes. A cast() that narrowed nothing failed strict typing, so this tag is a tree
that passes every check.
1.0.0 — 2026-08-29
Added
- A verified identity gets its own workspace, its own sessions, and a way out. The workspaces
root becomes that person's own directory, which narrows the switcher, the sidebar and every
session key at once. Storage keys on
(issuer, subject)and never on the address. - The sidebar names the signed-in person and offers Sign out when
signOutPathis set. - A personal workspace is seeded like any other, once. The credential store stays out.
- The container image is published:
ghcr.io/nanoinfraorg/nanoinfra. examples/auth/runs behind Caddy, with the three details that do not fail loudly written down.
Security
- Another person's workspace answers
403 that workspace is not yours; so does the shareddefault. Somebody else's session answers404rather than403, because a 403 says it exists.
Before 1.0.0
Thirty-four tags ending at v0.17.5, from before this repository took its current shape and mixed with upstream imports. That history lives in the release archive and on the releases page.