Skip to main content

Changelog

Releases before 1.0.0 are recorded in long form on the Release Archive, which is the only account of the 0.x line.

All notable changes to nanoinfra are recorded here.

The format follows Keep a Changelog 1.1.0, and this project adheres to Semantic Versioning.

A line names an effect a user can observe and carries a reference. The reasoning behind a change — what was measured, what was wrong first — lives in the issue or the proposal that line points at, not here.

Unreleased​

2.3.0 — 2026-09-08​

Added​

  • The composer says how much of the context window a thread is using, and what the last eight provider calls cost. Each bar splits the call's input into the three buckets billed separately, so a warm prompt and a cold one of the same size no longer look identical.

Fixed​

  • A streamed segment carries the cost of the provider call behind it. The value reached the channel's own tests but never a real turn: the loop wraps the delivery callback, and the hook offers a fact only to a callback that names it — so 51 stream_end records across 13 sessions arrived with no usage at all.

  • A socks:// proxy configured for the OpenAI-compatible or xAI provider now connects. httpx knows no such scheme, so on a host carrying that spelling every model call failed while its HTTP client was still being built.

  • A tool-call id issued by an earlier response is no longer replayed as a Responses input item id. After a restart, a model switch or any provider-state mismatch the endpoint could reject the whole request as an item that does not belong to this connection.

  • A Grok hosted search the stream abandoned no longer reads as a finished answer. The request is retried instead of answering a search question from nothing, and the search's activity row stops saying "searching" forever.

  • A relative working_dir resolves against the workspace. exec ran the command in whatever directory the gateway process was started in, and a CLI app refused the same path as "outside the configured workspace" — one input, two wrong answers in opposite directions.

  • Two MCP tools with non-ASCII names both stay registered and callable. Sanitization erased everything that told mcp_weather_获取天气 apart from its neighbour, so the second tool replaced the first while the connect log still counted both.

  • AGENTS.md, SOUL.md and USER.md are bounded in the system prompt, and say so when they are shown in part. They were embedded whole on every turn, and Dream is the writer of two of the three.

  • The idle-session summary cache is bounded. Entries were dropped only when a session was reopened, so a session archived and then abandoned held its slot for the life of the process.

  • A Codex chat now keeps one prompt cache key for its whole length, so the prefix it already paid for is billed as a cache read instead of being re-tokenised every turn. (#203)

  • Codex builds its TLS context once per process rather than once per request, so streaming no longer stutters while the CA bundle is read again before each call.

  • A background task that raises is now logged with its traceback and the name of the coroutine that failed. Memory consolidation and idle auto-compaction run through that path, so either could have been failing on every turn with nothing but asyncio's unattributed "Task exception was never retrieved" to show for it.

  • Fourteen log calls that passed exc_info=True now record the traceback. loguru takes keywords as format bindings, so the flag was accepted, unused and dropped — every one of those sites believed it was keeping an exception's traceback and was keeping none.

  • Six log messages that used printf %s now interpolate. The message printed a literal %s and the argument — an entry-point name, a failing tool's name — was dropped.

  • A provider backoff is visible on any channel that already takes progress updates. The event was published and then discarded before dispatch, so outside the CLI a two-minute Retry-After looked like a chat that had simply gone quiet.

  • A long Anthropic answer is no longer cut off at the stream idle timeout. The 90 s bound measured total generation instead of silence whenever no streaming callback was attached — which is every retry after a stall. (4de728a5)

  • Guidance queued in one chat is no longer sent into another. Opening a second chat that was idle read as the first chat finishing, so the waiting prompt went to whatever conversation was on screen; it now stays in its own chat until that chat's run ends.

  • One slow WebSocket client no longer delays the frames owed to every other client on the same chat, and a client that stops reading is disconnected rather than allowed to grow an unbounded outbound backlog in server memory.

  • An email whose hand-off to the agent fails stays unread and is delivered again on the next poll. It was marked \Seen and deduped during the fetch, so a failed hand-off was never retried and the mailbox reported it as handled.

  • A filtered email — self-sent, failing SPF/DKIM, or not allow-listed — is left unread. Marking it read made the mailbox's unread state stop reporting what the bot processed; postAction still decides what happens to it.

  • A degraded consolidation no longer loses the messages past the first 16,000 characters of its raw dump. When the summarising model errors or runs out of room, the batch is dumped to history as (part i/n) entries instead of one truncated entry — the compaction cursor advances past that batch either way, so the part that used to be cut off left the session with no copy anywhere. (#109)

  • Editing a cron job no longer cancels the job that is running. The scheduler's timer task is the task running the turn, and every edit re-armed that timer — so the agent's own cron tool, an operator toggling an automation in the WebUI, and a commissioning verdict each killed the turn mid-run and left the job to fire again. (upstream PR 5686)

  • ** in a find_files or grep glob now spans any number of directories, including none. The example both tools hand the model, tests/**/test_*.py, matched only paths with exactly one directory in between, and src/**/*.py matched nothing at all. (upstream PR 5692)

  • A consumer slower than the agent no longer loses an event from an SDK stream. Closing the stream evicted the oldest queued event to make room for its end marker, which in practice cost the final text_completed. (upstream PR 5635)

  • A backgrounded nanoinfra process now writes its print() output to its log. Redirected output is block-buffered, so a child that hung or was killed left an empty log file, exactly when it was being read to find out why. (upstream PR 5412)

  • The retry that drops an image and succeeds is counted. Only the failed attempt was recorded, so the tokens the answer actually cost were charged to nobody. (#176)

  • A model that raises instead of answering now fails over. An unauthenticated GitHub Copilot, or an endpoint that refuses the connection, skipped every configured fallback and was not retried.

  • A dropped or reset connection is retried. ConnectError, ReadError and RemoteProtocolError were reported as errors of no known kind, which made a transport failure look permanent.

  • An OpenAI server_error is retried instead of ending the turn on the first attempt.

  • A primary with an open circuit and no usable fallback returns the primary's own error, Retry-After included, and asks to be retried once the cooldown is over.

Changed​

  • nanoinfra webui starts even when model setup is incomplete, on every run and under --yes, naming the provider at fault. The second run used to refuse, which shut a half-configured install out of the Settings → Models screen that repairs it.

Removed​

  • The CLI Quick Start offer from nanoinfra webui. Provider and model setup is finished in WebUI Settings → Models.

  • An SDK run started with ephemeral=True writes nothing. The turn was persisted like any other — session file, mid-turn checkpoints and the cached session object — while only the session_turn_persisted event was withheld.

2.2.8 — 2026-09-07​

Changed​

  • mcp is capped below 1.30. That release changes three defaults at once — redirects are followed only within the endpoint's origin, idle Streamable HTTP sessions expire, and the OAuth client validates the authorization server's issuer — and the first two govern how this project talks to an MCP server over HTTP.

Fixed​

  • A command in the bwrap sandbox now runs nanoinfra's own Python. It resolved to the base interpreter instead, so a dependency this project declares and installs — openpyxl, and anything else a shell command imports — failed inside the sandbox while sitting installed. (#276)
  • A redirect to /dev/null no longer counts as a workspace-bypass attempt. Three of them in one turn — which any non-trivial shell command reaches — told the agent it had hit a hard policy boundary about a path it did not want, and the turn derailed. (#282)

2.2.7 — 2026-09-06​

Added​

  • Metrics → Calls filters by source and actor. Both were already on every row and in the payload with no control to filter on them, so finding what one approver authorised, or what one channel ran, meant reading the table by eye. (#274)

2.2.6 — 2026-09-06​

Changed​

  • Abilities in the sidebar starts closed, and opens only when you click it. It opened by default, and it also reopened itself whenever the active page was one of its two members — so a visit to Skills undid the collapse. Both are gone: absent means closed, and the heading is the only thing that opens it. Infrastructure already started closed and keeps its own reopen behaviour, which was not part of the ask. (#253)
  • Both rail groups remember whether you closed them. They held that in local state, so the choice lasted until the next reload — and the comment on that state called collapsing "the operator's choice, not the default", which a choice that does not survive a reload is not. They now use collapsed_groups, the map the sidebar already round-trips for chat project groups, under nav:-prefixed keys so they cannot collide with a project of the same name. (#253)

2.2.5 — 2026-09-06​

Fixed​

  • A row's detail in Metrics → Calls opens under that row instead of after the whole table. Opening row 1 of 100 put its fields below row 100, so reading a call meant scrolling the page away from the row that was clicked. The chevron on the left already promised an inline disclosure, and a disclosure that opens a hundred rows away is a broken affordance rather than a layout preference. Thirteen fields also lay out in three columns on a wide screen now, so the rows below do not travel far. (#274)

2.2.4 — 2026-09-06​

Added​

  • The Live tab draws its numbers as well as printing them. Each gauge grows a sparkline of the last three minutes, kept in the tab — nothing stores gauge history, because the gauges are sampled when read, and the panel says so rather than implying it holds yesterday. A gauge that could not be read gets its dash and no plot: an empty plot area reads as flat at zero. (#274)
  • Context used and its limit are one meter instead of two tiles. They were never two facts, and reading 428K against 1,048,576 is arithmetic the panel should do. The fill carries severity past 75% and 90%, with the percentage always spelled out — a status colour never carries meaning alone. An absent or zero limit reads as unknown rather than as 0%. (#274)
  • Two charts under the gauges: calls per minute, derived from the difference between successive reads of the cumulative counters the way rate() does, and a latency histogram over the nine buckets. An empty latency bucket keeps its row, because a band with no calls is information. GET /api/webui/metrics/counters serves both. (#274)

Fixed​

  • The counters charts cannot take the Live tab down. They render inside it, so trusting the route's shape meant an older gateway — or any unexpected answer — unmounted every gauge above them. The same failure the scale row had one release earlier. (#274)

2.2.3 — 2026-09-05​

Added​

  • An Approvals tab in Metrics: how many actions the gate held for a person, how many they answered, how many they refused, how many expired, and the median time to answer — the number that says whether the gate is working rather than merely running. A person's refusal is counted apart from a policy refusal nobody was asked about, because merging them overstated the approver's denials sixteen-fold on the deployment this was measured against. An ask that was neither answered nor expired is named on its own: nothing ran and nothing said why. (#274)
  • A scale row at the top of Metrics → Usage: servers, skills, agents, MCP servers and connectors, in one request. Every one of these existed and was scattered across five settings pages. A count that cannot be read shows — and names itself rather than reading as zero. (#274)
  • /metrics exports counters and a latency histogram, so a Prometheus install can ask for a rate and a quantile rather than only a level: nanoinfra_llm_calls_total, nanoinfra_llm_tokens_total (by input, output, cache read and cache write), nanoinfra_tool_calls_total, and nanoinfra_llm_duration_ms over nine buckets. Accumulated in memory for the life of the process, because a count over a table with a purge is not monotonic and a rate() over a falling counter is nonsense. (#274)
  • Two process-health gauges, nanoinfra_rss_bytes and nanoinfra_event_loop_lag_ms. A rising lag is a blocked loop, and neither number has an event to be driven by, so both are sampled at read. (#274)

Fixed​

  • The gate audit viewer no longer offers deny as a decision filter. Nothing writes it: the name lives on as the outcome enum and as the operator socket's wire verb, and the log records that outcome as denied so it speaks the operator's vocabulary. The filter could only ever match zero records. (#274)

2.2.2 — 2026-09-05​

Fixed​

  • The socket directories are setgid for real this time, verified in the built image rather than reasoned about. v2.2.1 set the mode before the chown in prepare, which was right and not enough: the post-bind block re-applies the mode, and by then the directory already carries the helper's group, so its chmod 2710 dropped the bit again and returned success. All five directories were still at 710 on 2.2.1. One set_socket_dir_mode helper now owns that mode and takes the directory back to root before setting it, which is the only order that works on a fresh directory, on a re-apply, and on the 0700 one the Python side creates.

2.2.1 — 2026-09-05​

Fixed​

  • Every helper socket directory is actually setgid now, so a socket keeps its shared group when the helper rebinds it. chmod 2710 ran after chown, and without CAP_FSETID — which the published compose file does not grant — that silently drops the setgid bit and returns success, so all five directories sat at 710 while the code's own comments described 2710. The operator socket is the one that paid: the executor deliberately does not join nanoinfra-op, so the inherited group is the only mechanism it has, and its absence left the racy root chown as the only thing setting it — plus a [Errno 1] Operation not permitted on every boot.
  • apply_socket_group sets the mode even when it cannot set the group. The two were in one try, so a refused chown skipped the chmod — and the mode is the half that grants a peer its write bit.

2.2.0 — 2026-09-05​

Added​

  • A Metrics destination in the rail, with three tabs. Usage shows spend per model over a chosen window of 7, 30, 90 or 365 days, with cache writes, truncated answers, time to first token, wall clock and a per-model cost — five of which the store has recorded since 2.0.0 and none of which reached a pixel. Live shows seven point-in-time gauges, starting with how many suspended actions are waiting for a person. Calls reads the tool_calls table, which had a writer, a pruner and a purge log and no reader at all. (#235, #232)
  • pricing in config gives a model four rates in USD per million tokens — input, output, cache read and cache write — keyed "<provider>/<model>". Until one is set the Usage tab says no prices are configured rather than showing a spend of $0.00. (#235)
  • gateway.metricsEnabled serves the same gauges in Prometheus text format at /metrics on the gateway's own port. Off by default. gateway.metricsToken sets the bearer token a scrape must present; with no token only a loopback bind is served, so enabling metrics on a port a reverse proxy fronts does not publish them. (#235)
  • The Usage tab lists why calls failed, per error kind, status code and provider. The failure count was already shown; the reason behind it was recorded and never read. (#235)
  • The Usage tab breaks the window down by what started the turns — chat, API, automations, memory, system. source was aggregated per day and readable only inside one heatmap cell's tooltip, so "what does automation cost me this month" meant opening thirty tooltips. (#235)
  • A model row in the Usage tab expands to the four measurements the columns are checked against: how many calls the provider reported against how many were tokenized locally, output tokens per second of generation, measured against reported output, and the streamed split that time to first token is averaged over. (#235)

Fixed​

  • A cached token is no longer billed twice. prompt_tokens is the logical input and includes the cached halves, so charging it at the input rate and the cached count at the cache rate over-stated a warm cache badly — 5.7× on published Kimi K3 rates for a 90%-cached prompt. The three input buckets are now disjoint. (#235)
  • A usage row names the provider that was configured rather than the class that made the call. OpenAICompatProvider serves every OpenAI-compatible API, so Moonshot, DeepSeek, Groq, OpenRouter and forty others were all recorded as openaicompat — which left "which provider is expensive" unanswerable and meant a rate set for moonshot/kimi-k3 could never match its own rows. Single-provider backends were wrong too: openai_codex recorded openaicodex. (#235)
  • Both spellings of the pricing rate keys are accepted (inputPerMtok and inputPerMTok), and an entry that states no rate nanoinfra recognises reads as unpriced rather than as $0.00. A mistyped key used to produce a confident zero over a month of real spend. (#235)

Changed​

  • Model rates are set where the model is: Settings → Models → Pricing, with a live figure showing what the recorded window would have cost at the rates typed, a warning when the model reads cached tokens and the cache rate is still zero, and a note when two configurations name the same model and therefore share one bill. Settings → Providers → Default pricing sets rates for a whole provider, including a free checkbox — the one-edit answer for a local fleet. (#235)
  • The Settings overview no longer opens with the token summary and the heatmap. They are the Usage tab now, and the overview keeps one row that leads there — which also ends the five-second re-aggregation of llm_calls that a settings page left open used to run. (#235)

2.1.0 — 2026-09-04​

Added​

  • A tool group, MCP server or connector can be set to attach: "search": its schemas leave every prompt and the model loads them itself by calling the new tool_search tool, with one shared pointer in place of a per-item advertised line. mention stays the user-driven deferral; both respect the acting agent's ceiling, which neither can widen past.

Fixed​

  • A tag no longer publishes to PyPI or GHCR before the Test Suite has finished on that commit. v2.0.0 shipped a wheel that could not be imported on the minimum supported Python because the publish jobs are quicker than the tests and nothing connected them. (#270)

Changed​

  • CI builds the image with Buildx and the Actions cache, the way the publish workflow already did. (#271)

2.0.4 — 2026-09-04​

Fixed​

  • @agent:<name> in a message answers as that agent. The composer offered the token, completed it, and then ignored it: the turn ran as the deployment default, so naming an agent looked like it did nothing. (#269)

2.0.3 — 2026-09-04​

Fixed​

  • The WebUI inbox is no longer warned about as an unreachable chat channel. A deployment whose approver sits there was told on every poll that a suspended action reaches nobody, while the inbox had the request. (#267)

2.0.2 — 2026-09-04​

Fixed​

  • The Agents destination is in the sidebar whether or not the deployment names an agent. It was gated on the roster, so on a fresh install nothing led to the page that configures the agent answering every turn. (#266)

Changed​

  • Abilities is no longer gated on the roster either, and it opens expanded — one sidebar shape whatever the deployment holds, with Apps and Skills still one click away. (#253)

2.0.1 — 2026-09-04​

Fixed​

  • nanoinfra imports on Python 3.11 again. A dataclass field on RequestContext defaulted to a mappingproxy, which 3.11 refuses for any default whose type is unhashable — so 2.0.0 failed at import on the minimum supported version. (#266)

2.0.0 — 2026-09-04​

Added​

  • One approval can become a standing grant. The grant is derived from the payload the executor actually rendered, defaults to expiring, and asks once more before it never does. (#217)
  • Every server keeps notes the agent and the operator both write, and a turn that names a server reads them. A note does not expire; one that disagrees with what you see is evidence the infrastructure changed. (#222)
  • A queryable record of every tool call: which tool, by whom, what the gate decided, and how it ended. It stores the address of the arguments in the session history, never the arguments. (#231)
  • The agent can search your own documents and answer with citations. Drop files in <workspace>/knowledge/; nothing reaches a prompt on a turn that does not ask. (#237)
  • A deployment can name more than one agent, each with its own model, tools, skills and instructions. An empty agents.named is exactly the single agent every deployment has today. (#247)
  • An agent can hand one task to a peer and wait for its answer. Membership in agents.named[x].delegates is the grant, and delegation is one level deep. (#250)
  • A delegated action records the human who asked, the agent that delegated and the peer that acted, so a reader can answer "who authorised this" without opening a second file. (#251)
  • Every assistant turn says which agent answered it, beside the turn's cost. (#248)
  • The composer offers the agents a message may ask for, as @agent:<name>. The token stays in the text, because it is a preference the answering agent reads and not an invocation. (#255)
  • A turn that delegates shows its plan as one object in the thread — a row per peer with its outcome and its own cost — and a reload shows the same plan. (#252)
  • An Agents destination lists the agents a deployment names, and an Abilities grouping collects Apps and Skills in the menu. (#253)
  • Each agent's Prompt tab shows the prompt's sections with the permission on each, and the addendum that appends after them. (#256)
  • An automation can name the agent it runs as, and that agent is the ceiling: its tool groups cap the turn and its skills bound the job's own picker. (#257)
  • The approvals inbox names the agent that will act, and the agent that delegated to it. (#258)
  • Agents are created, edited and deleted in the browser. Each one gets a page with tabs — model, tools, skills, delegates, prompt — and every binding is picked from what this deployment has. (#262)
  • The deployment's own agent is one more agent: same page, same tabs, editable down to its skills, MCP servers and delegates. It still cannot be deleted. (#265, #266)
  • The Prompt tab edits the prompt. The three sections that are prose can be replaced, each with the text in force shown and what replacing it costs said before you do. (#256)
  • Settings → Prompts reads and writes the two prompts that run unattended, dream and evaluator. (#264)
  • An agent's tool groups, skills, MCP servers and connectors now narrow every turn it answers, not only a scheduled one — so choosing an agent is how a conversation stops paying for every server and skill installed. (#266)

Changed​

  • Where a deployment names agents, the composer chooses an agent instead of a model — the model belongs to the agent. A deployment that names none keeps its model selector unchanged. (#254)
  • The prompt's safety notes are their own fixed section, so replacing the runtime section can no longer delete them. (#256)
  • A replaced prompt section is still named in the prompt manifest and marked as overridden — a record that hid a replacement would make two different prompts look identical. (#256)
  • The executor protocol is version 7, carrying the delegation chain. A deployment running the executor as a separate process must restart it alongside the gateway. (#251)
  • Nothing ships a model. agents.defaults.model was anthropic/claude-opus-4-5, so an unconfigured deployment looked configured on a provider it had no credential for. The first model configuration a deployment adds is now the primary one, and it answers until something else is chosen. (#266)
  • An agent's empty binding list means none of them, not all of them. Declaring nothing is what means everything, and the two are now separate answers everywhere: config, the tool filter and the picker. (#266)
  • A persona can write {{ agent_name }}, {{ agent_role }} or {{ agent_description }} in SOUL.md and each agent fills it in. Any other {{ }} survives verbatim. (#265)

Removed​

  • The config warning about a dead agents.defaults.model. The turn fell back to the primary preset either way, so the line reported a difference nothing acts on — and the field ships empty now. (#205)

Fixed​

  • A system job kept its state across a restart. dream and evaluator were re-registered as brand new on every boot, so on a deployment that restarts they never ran. (#263)
  • An agent's addendum and its replaced prompt sections reach the model. Both were stored, shown, editable — and never passed to the prompt builder by its only call site. (#265)
  • A tool refused for want of a secret says when the encryption key is the wrong one, instead of failing silently. (#217)
  • A trigger whose agent declared no tool groups is no longer capped to the ungrouped tools. (#266)
  • The default agent has a card on a fresh install. It hid until the deployment named an agent, so the agent that answers every turn was reachable only after naming one you did not want. (#266)
  • A deployment with no model at all says so at boot and on the settings page, instead of failing a turn with No provider is configured for model ''. (#266)

Security​

  • A delegated turn can no longer reach a capability class the turn that spawned it did not hold, and a capped turn hands that ceiling to anything it spawns. (#251)

1.9.1 — 2026-09-03​

Removed​

  • The sidebar no longer lists conversations from other channels, reverting 1.9.0. Sessions that never held a conversation — an empty WhatsApp session, cli:direct — showed up as New topic rows and buried the chats, which is worse than the problem it set out to fix. (#216)

1.9.0 — 2026-09-03​

Added​

  • A conversation held over the API — or from a chat channel, or by an automation — appears in the WebUI sidebar with its channel badged, and opens as a readable thread built from the session history. Read-only: the composer still refuses a session it does not own, and delete, file-preview and automations stay closed to other channels. (#216)

1.8.0 — 2026-09-03​

Added​

  • api.enabled lets the gateway serve /v1 on api.port itself: one process and one agent loop instead of two, sharing the MCP and connector hosts it already booted. Default false, so upgrading opens no port. nanoinfra serve is unchanged. (#214)
  • The API server logs one line per request — method, path, status, duration, and why a request was refused — and its own logger stays audible without --verbose. Never the body or the Authorization header. (#215)

1.7.4 — 2026-09-03​

Fixed​

  • nanoinfra serve drains the agent's outbound event bus, so a turn that emits more than a thousand progress or stream events finishes instead of stalling mid-flight and leaving its HTTP request unanswered. gateway and nanoinfra agent already drained; the API server did not. (#211)

1.7.3 — 2026-09-02​

Fixed​

  • Both API routes accept a client that resends its own transcript: what follows the last assistant message is the turn, and earlier messages are dropped as the client's copy of a history the server already keeps. A system message beside the prompt is joined into the turn rather than refused, so a Responses or Chat Completions client no longer gets a 400 on its first request. (#211)

1.7.2 — 2026-09-02​

Added​

  • The prompt breakdown says which request it describes when a turn made more than one, and names the largest request the turn reached. (#208)
  • An activity cluster names what the provider calls inside it cost — input, cache share and output — beside the duration it already showed. The cache share counts only the calls that reported one. (#208)
  • tools.groups declares groups of built-in tools with the always | mention modes an MCP server and a connector already had, so @diagrams can load 2,438 tokens of diagram schemas only for a turn that asks for them. diagrams and servers are predefined; both default to always. (#210)
  • POST /v1/responses answers the same agent over the Responses wire, so a client that defaults to that protocol no longer gets a 404. The caller's tools and instructions are ignored: the agent runs its own tools behind the capability gate. (#211)
  • agents.defaults.midTurnMessages decides what happens to a message that arrives while a turn is running. The new default, queue, gives it a turn of its own with its own answer; inject keeps the previous behaviour of folding it into the turn in flight. (#209)

Changed​

  • A message sent while the agent is still working now gets its own turn and its own reply, instead of being folded into the turn already running and answered only as part of it. (#209)

Fixed​

  • Each activity cluster reports its own duration instead of the whole turn's, so eight consecutive steps no longer all read the same figure. (#208)

1.7.1 — 2026-09-02​

Added​

  • /compact archives a session's history on request instead of waiting for an idle timer or a budget threshold, and reports how many messages it archived and how many stay raw. (#212)

1.7.0 — 2026-09-01​

Added​

  • OpenAI requests carry prompt_cache_key, one key per chat, so its automatic prefix cache is reached instead of every session landing in the bucket its first 256 tokens hash to. Opt-in per provider: it is a body field, and a provider that rejects an unknown one answers 400.
  • google_calendar_freebusy answers whether a calendar is busy in a range — busy blocks only, no event detail — so availability and slot-finding no longer mean reading every event. A connector operation can now declare read_via_post for a read whose query needs a POST body; the invariant that a POST is a write otherwise still holds.
  • google_calendar_update_event changes an event by id. A partial PATCH, so an omitted field keeps its value — sending only start/end moves an event and keeps the rest. mutate.remote.
  • google_calendar_delete_event removes an event by id. It carries mutate.remote, the same class as creating one, so it asks a person in an interactive turn and needs a standing grant to run unattended.

1.6.1 — 2026-09-01​

Fixed​

  • Every turn of one chat reaches the same xAI server, so its prompt prefix can be cached. x-grok-conv-id was a fresh UUID per request, which announced each call as a new conversation.

Added​

  • CHANGELOG.md, in Keep a Changelog form, covering every version since 1.0.0. A release body is now this file's section for that version.
  • A tag whose version has no changelog section fails the release workflow.

1.6.0 — 2026-09-01​

Added​

  • The prompt breakdown's tool count opens the list of tools behind it, each with its size, largest first. (#203)

Changed​

  • A tool row in the breakdown reads as google-calendar with a connector tag, rather than as the raw connector:google-calendar. (#203)

1.5.2 — 2026-09-01​

Fixed​

  • Every reference prefix keeps a row in the mention menu. The palette reserved two, so a deployment with a data connector — server:, diagram:, calendar: — lost one. (#204)

1.5.1 — 2026-09-01​

No user-visible changes. A route test pinned the arguments the marketplace search is called with, and 1.5.0 added one.

1.5.0 — 2026-09-01​

Requires skills-server v0.4.0.

Added​

  • Apps → Connectors browses the catalog and installs from it. Each row names every operation with its capability class, the hosts a token could reach, and the scopes it would carry, before the install button. (#207)
  • An install reports what is still missing: a connector has no credential and no entry in connectors.active until an operator adds them. (#207)

Changed​

  • The marketplace reads the catalog's kind and installs each package where its own subsystem reads it — skills/, plugins/, connector-packages/. The kind is asked of the catalog and never taken from the caller. (#207)

Fixed​

  • A connector already on disk read as not installed, because installed was asked of the skills loader for every kind. (#207)

Security​

  • Workspace containment runs before anything walks the install destination. A symlinked skills/ was read through before the boundary check. (#207)

1.4.0 — 2026-09-01​

Added​

  • attach: "always" | "mention" on a data connector. A mention connector stays active with its credential and grants, the prompt carries one line naming it, and its operations load only for a turn that names it — @<name>, or a @<kind>:<id> object of one of its kinds. (#204)
  • An automation declares connectors: [...], since an unattended turn types no @. (#204)

Fixed​

  • The composer palette silently dropped candidates whose kind was missing from its group list. (#204)
  • A paused MCP server named in message text was still tokenised and still sent. (#206)

1.3.2 — 2026-09-01​

Fixed​

  • The connector mention prefix reads the cached object listing, so @calendar: is in the menu on the first keystroke instead of waiting on a live API call that took seconds and failed silently. (#204)

1.3.1 — 2026-09-01​

Fixed​

  • A paused MCP server is no longer offered in the @ palette. Picking one read as an attachment on a server nanoinfra never connected. (#206)

1.3.0 — 2026-09-01​

Added​

  • attach on an MCP server. A mention server is advertised in one line — name, tool count, how to attach — roughly 50 tokens against a couple of thousand, with its schemas sent only to a turn that names it. (#204)
  • An automation declares mcpPresets: [...]. (#204)
  • Config that is ignored says so, at boot in the log and on the Models page: a dead agents.defaults.model beside a modelPreset that overrides it, and a selected preset whose provider has no credential. (#205)
  • The Apps page's tool count opens the tool list.

Changed​

  • The Apps page opens on Ready rather than the CLI tab.

Fixed​

  • A paused row counts its saved allowlist instead of reporting 0 tools for a server holding fifteen.

1.2.3 — 2026-09-01​

Fixed​

  • Pausing an MCP server stops it connecting. AgentLoop.from_config seeded the live map from the unfiltered config, so a paused server reconnected on the next restart. (#206)
  • A plugin-declared MCP server reached the loop only after a hot reload. (#140)

1.2.2 — 2026-09-01​

Changed​

  • The MCP pause is a switch on the row; the row's checkmark becomes an explicit ⋯ menu, and the second line names what the server costs instead of naming its transport. (#206)

1.2.1 — 2026-09-01​

Added​

  • An MCP server can be paused without losing its configuration — enabled: false. Its command, arguments, environment, headers and enabledTools list stay; its schemas leave every prompt. (#206)

1.2.0 — 2026-08-31​

Added​

  • A per-turn prompt breakdown, collapsed under each turn: where the input tokens went, by section and by tool source, largest first. Recorded while the prompt is assembled, because the attribution is gone by the time a request reaches a provider. Names and sizes only, never content. (#203)

1.1.4 — 2026-08-31​

Added​

  • Settings leads with the token numbers: 30-day tokens, calls, failures, today, the measured share, the peak day, and a per-model breakdown with time to first token. (#176, #177)

Fixed​

  • A migrated day topped the per-model breakdown with a day's worth of tokens against 19 "calls". It counts in the day totals and not in a breakdown it cannot answer.
  • Every fallback row was labelled fallback rather than the provider that actually answered.

1.1.3 — 2026-08-31​

Fixed​

  • A connector's stale last_error can be retracted. A successful test or a successful listing clears it; every success had been passing an empty value, which the merge ignores by design.

1.1.2 — 2026-08-31​

Fixed​

  • Every connector call failed with [Errno 13] Permission denied on any deployment that had activated one. Marketplace package discovery pointed at the connector state directory, which the executor cannot traverse, and Path.is_file() propagates a permission error rather than answering False.
  • The connector row in Apps spans both grid columns, so its name is readable and its capability classes read as chips rather than as a vertical wall.

Changed​

  • Connector packages get their own root, <workspace>/connector-packages/. A package is read by three accounts; a state file is written by one.

1.1.1 — 2026-08-31​

Added​

  • A connector against a public API runs with no credential — credential.kind: "none" activates with no binding, mints nothing and sends no Authorization header. It is still gated per operation.
  • examples/connectors/hello-world/: one read against a public API, no credential, no setup. Ships in the sdist, because a pip user has no checkout to copy from.

1.1.0 — 2026-08-31​

Added​

  • Every finished turn shows its cost in the footer — tokens in, out, cache share and latency — persisted beside the latency so a reloaded thread shows what the live turn showed.
  • One content-free row per provider attempt in ~/.nanoinfra/llm-usage.sqlite3, retained 400 days. No prompts, no responses, no reasoning, no tool payloads, no session keys; a test asserts the schema has no column content could go in.
  • A connector can arrive from the catalog as a declarative connector.json with nothing importable, and its calls run in a fourth confined process (nanoinfra-connector, outbound 443 only) whose group holds the executor and not the agent.

Changed​

  • LLMResponse.usage and both hook contexts carry LLMUsage | None instead of a dict. to_turn_dict() is the projection the OpenAI-compatible API, the SDK and the WebUI read.
  • ~/.nanoinfra/webui/token-usage.json is migrated at first start and renamed, not deleted.

Fixed​

  • Three usage-accounting defects the typed contract exposed: a cache count present on one call and absent on another summed as though both were measured; two calls carrying 40k of context each read as a turn carrying 80k; the over-budget finalisation path mutated the accumulator through its argument.
  • Three call sites reached a provider without passing through the agent loop — WebUI title generation, the evaluator, Dream consolidation — so their tokens were charged by the provider and counted by nothing.

Security​

  • A connector credential is bound to the hosts it may address, checked at activation with both hosts named. A package declaring Google scopes and its own baseUrl would otherwise receive a live Google token and a process that forwards it.

1.0.6 — 2026-08-31​

Fixed​

  • A secret record written through the executor took the writer's primary group, so a group-readable file in the wrong group answered EACCES to the gateway: a create succeeded and the next listing raised. The store sets the group explicitly when the directory shares one.

1.0.5 — 2026-08-31​

Added​

  • The Secrets page writes on a container deployment. Protocol version 6 carries a third request kind: the gateway encrypts because it holds the key, and the executor writes ciphertext it cannot read. Three verbs — create, update, delete — and no read.
  • Connector consent completes in the browser, returning to this deployment's own origin with the credential and the activation written. A consent that fails writes nothing.
  • Reload connectors reconciles the running registry with config, without a restart.

Fixed​

  • The no-tools request path had no timeout while the tool-carrying path wrapped the same call, so a stall held the per-session lock forever.
  • A truncated consolidation was accepted as history, taking whatever it had not reached with it.
  • find_files and grep bounded their results and nothing about their scan, and walked inside an async execute — one call on a slow filesystem held the event loop for every session. Both carry an entry budget and a wall clock, and a partial answer says so.
  • A cron job froze the workspace path one person's turn resolved, naming their identity directory, into jobs.json.

Security​

  • A Slack file download followed any URL the workspace handed it, redirects included, with no SSRF validation and no DNS pinning.

1.0.4 — 2026-08-30​

Added​

  • Data connectors. A connector reaches one data source with a capability class per operation, so reading a calendar and writing to it are two decisions — where an MCP tool declares nothing and every one of them resolves to the fail-closed mutate.remote.
  • google-calendar ships first, with three read operations and one mutate.remote.
  • The call runs in the executor: protocol version 5 carries a second request kind, and the method, path, class and scopes come from the installed manifest — so a frame cannot describe a call the package never declared, and the agent process holds no token.
  • A standing grant can name a connector, so an unattended write is expressible rather than permanently denied.

1.0.3 — 2026-08-29​

Fixed​

  • The test that pinned the previous dependency policy now states the current one. v1.0.2 moved four packages into the base install and left an assertion saying the opposite, so the release shipped and CI failed.

1.0.2 — 2026-08-29​

Changed​

  • asyncssh, ansible-runner, boto3 and aiohttp are base dependencies. The container image installs no extras, so the published image came up as an agent for infrastructure that could not reach infrastructure. nanoinfra[servers] and nanoinfra[api] still resolve.

1.0.1 — 2026-08-29​

No behaviour changes. A cast() that narrowed nothing failed strict typing, so this tag is a tree that passes every check.

1.0.0 — 2026-08-29​

Added​

  • A verified identity gets its own workspace, its own sessions, and a way out. The workspaces root becomes that person's own directory, which narrows the switcher, the sidebar and every session key at once. Storage keys on (issuer, subject) and never on the address.
  • The sidebar names the signed-in person and offers Sign out when signOutPath is set.
  • A personal workspace is seeded like any other, once. The credential store stays out.
  • The container image is published: ghcr.io/nanoinfraorg/nanoinfra.
  • examples/auth/ runs behind Caddy, with the three details that do not fail loudly written down.

Security​

  • Another person's workspace answers 403 that workspace is not yours; so does the shared default. Somebody else's session answers 404 rather than 403, because a 403 says it exists.

Before 1.0.0​

Thirty-four tags ending at v0.17.5, from before this repository took its current shape and mixed with upstream imports. That history lives in the release archive and on the releases page.