Automations
Automations are agent turns that run later in a linked topic. Use them when nanoinfra should do work without someone actively typing: reminders, recurring checks, nightly summaries, CI follow-ups, local script reports, or webhook-driven events.
Create automations from the chat channel or WebUI topic where the result should appear. That lets nanoinfra keep the right session history, workspace, and reply target.
Is This an Automation at All?
Answer this before picking a type. A sustained goal is the other shape of long-running work, and it is not an automation:
| Sustained goal | Automation | |
|---|---|---|
| Started with | /goal <goal> in the chat | see the table below |
| Runs | now, across turns, until the goal is met | when its schedule or trigger says |
| Between runs | holds the session's context, and remembers what it already tried | starts clean, and keeps only the automation_state it wrote on purpose |
| Use it for | work that builds on itself | a check that should answer the same whatever ran yesterday |
So a review that needs to know what it already reviewed is a goal. A morning summary is an automation. See Long-Running AI Agent for both, side by side.
Choose an Automation Type
| Type | Starts from | Best for | Created with |
|---|---|---|---|
| Scheduled automation | Time, interval, or cron expression | Recurring reminders, scheduled summaries, one-time future tasks | Ask nanoinfra in the target topic to schedule it with the cron tool |
| Local trigger | A local nanoinfra trigger ... command | CI jobs, webhooks, shell scripts, generated reports | /trigger <name> in the target topic |
| Heartbeat | Protected system schedule | Quiet recurring checks that should only report useful results | Edit <workspace>/HEARTBEAT.md |
The two user-created automation types are scheduled automations and local triggers. Heartbeat uses the same background service but is system-managed and protected from normal automation edits.
Before You Create One
Keep nanoinfra gateway running. The gateway owns background delivery for chat
apps, WebUI topics, scheduled automations, local triggers, heartbeat, and
Dream jobs.
Use the same workspace and config for the gateway and any process that sends
local trigger messages. If you run multiple nanoinfra instances, pass the matching
--config or --workspace option to nanoinfra trigger.
Create each automation from the target topic. An automation without a linked topic cannot be enabled or run from the WebUI because nanoinfra would not know where to deliver the turn.
Creating One Rehearses It
Creating a scheduled automation runs it once, immediately, with every gated tool forced to preview. Nothing executes. The rehearsal resolves the action and asks the capability gate what a scheduled run would meet. Then it tells you, in the chat where you created it.
That exists because an automation is prose, and its commands are not in its text — the model composes them when it runs. So the only thing that can answer "what permission will this need?" is a run. Doing that run while you are still here beats learning the answer from a refusal at 03:00.
Three outcomes:
- Rehearsed clean. Every action it would take is already permitted. It runs on its schedule.
- Would be refused. It is saved disabled, with the finding attached and the standing grant that would permit it. Your work is not lost, and no automation sits enabled while certain to refuse unattended.
- The rehearsal did not finish. Nothing was learned — a provider outage, for instance. The automation is left alone, because a fault of the rehearsal is not a verdict about the automation.
In the WebUI, the automation's detail pane carries the finding and two actions.
Rehearse runs it again. Grant it writes the proposed standing grant.
Read what a grant covers before you press it: that command on those hosts in
any unattended turn, not this automation alone. See
capability-gates.md#commissioning-learn-the-grant-before-the-schedule-does.
A rehearsal costs a model turn, so editing an automation only re-runs it when the message, references or skills change. Renaming it or moving its schedule does not change which commands it runs.
A local trigger is not rehearsed, and says so at creation. Its content arrives from whoever fires it, so no rehearsal can learn which commands it will run. Rehearsing an invented message would propose a grant for a command the real firing may never use. Fire it once while you are watching instead.
Scheduled Automations
Scheduled automations are created by the agent's cron tool. In practice, ask
nanoinfra from the target chat or WebUI topic:
Every weekday at 9am, check open pull requests and summarize blockers here.
or:
Tomorrow at 4pm, remind me to send the release notes.
The cron tool supports interval schedules, cron expressions, and one-time
scheduled tasks. Cron expressions can include an IANA timezone such as
America/Vancouver. Otherwise nanoinfra uses the runtime default timezone.
Scheduled automations normally deliver the result back to the session where they were created. Use them for work that should run on a predictable schedule and report each run.
For background checks that should stay quiet unless there is something useful to report, use heartbeat instead of a user-created scheduled automation.
Local Triggers
Local triggers let a local script or external service send a message into a specific nanoinfra session later.
Create the trigger from the chat or WebUI topic where future messages should arrive:
/trigger PR review
nanoinfra replies with a trigger ID and a command shaped like:
nanoinfra trigger trg_8K4P2Q9X "Review PR #4502"
Replace the quoted text with the message nanoinfra should receive. For generated or longer content, pipe stdin:
generate-report | nanoinfra trigger trg_8K4P2Q9X
For multiple instances, use the same config or workspace selector as the gateway:
nanoinfra trigger --config ./bot-a/config.json trg_8K4P2Q9X "Nightly report"
nanoinfra trigger --workspace ./bot-a/workspace trg_8K4P2Q9X "Nightly report"
Firing a trigger over HTTP
Anything that can reach the gateway can also fire a trigger. That lets a monitor, a CI job, a backup, or a provider webhook start an automation without shell access on the gateway host.
Each trigger carries its own key. Issue one from the WebUI Automations view, or over the API with an operator token:
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
http://127.0.0.1:8765/api/webui/automations/trg_8K4P2Q9X/key
The response is the only place that key ever appears. Only its digest is stored, so nothing can read it back — if it is lost, issue another, which is also how rotation works. Issuing again invalidates the previous key.
Then fire the trigger with that key:
curl -H "Authorization: Bearer $TRIGGER_KEY" \
-H "X-Nanoinfra-Trigger-Message: Review PR #4502" \
-H "X-Nanoinfra-Trigger-Idempotency-Key: pr-4502" \
http://127.0.0.1:8765/api/triggers/trg_8K4P2Q9X/fire
What that key does and does not authorise:
| Fires | exactly that one trigger |
| Cannot | read sessions, list triggers, answer approvals, read secrets, issue keys, replay deliveries |
| Cannot | enumerate — a wrong trigger id and a wrong key return the same 401, so the endpoint is not a directory |
The message travels in a header rather than a body, because the gateway's transport exposes no body. It is not in the query string either. Every proxy the request passes logs a query string, and this content becomes part of a prompt. It is capped at 6000 bytes, and refused rather than truncated. A caller told its payload was rejected can adapt. One whose alert was silently cut in half cannot.
X-Nanoinfra-Trigger-Idempotency-Key is optional and recommended. A caller
whose POST times out and retries is the normal case. With the same key, the
second call answers 200 with {"queued": false, "duplicate": true} rather
than firing again. One trigger accepts at most one fire per second. Over that
answers 429 rather than silently queueing.
Revoke a key when the caller is retired:
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
http://127.0.0.1:8765/api/webui/automations/trg_8K4P2Q9X/key/revoke
The HTTP path only queues the delivery — it does not run the turn. So it gets
the same durable queue, retries, dead-lettering and audit records as
nanoinfra trigger, because it is the same queue.
Heartbeat
Heartbeat is for recurring workspace checks that should usually stay quiet. It
reads <workspace>/HEARTBEAT.md, executes active tasks, and sends only useful or
actionable results to the most recently active chat target.
Use heartbeat for checks such as "watch this repo for important failures" or "periodically inspect this workspace and only tell me when action is needed." Use a scheduled automation instead when every run should produce a visible reminder or report.
Heartbeat is enabled by default when nanoinfra gateway starts. Configure it in
configuration.md#gateway-heartbeat.
Manage Automations
Use the WebUI Automations view to:
- filter by all, active, paused, needs-attention, or system jobs.
- search by task name, message, trigger command, linked topic, schedule, or status.
- sort by next run, last run, updated time, or name.
- run scheduled automations now.
- pause or resume, rename, or delete user-created automations.
- copy the CLI command for local triggers.
- inspect protected system automations without changing them.
Local triggers do not have a WebUI "Run now" action because each run needs a
message. Copy the nanoinfra trigger ... command from the WebUI and replace
"message" with the content that should be delivered.
The detail pane also shows Recent runs and Remembered state. Recent runs holds the last five, with status, duration, any error, and why each one ran. Remembered state has a reset.
Fields That Used To Live in the Prompt
An automation's message is prose, and that is the design: the agent works out the steps rather than you wiring them. But some things a prompt holds badly, and those are now fields in the editor's Advanced group. It is collapsed by default, because every field in it already holds the behaviour automations had before.
Remembered state
An automation that needs to remember something between runs used to have to
maintain a file by hand. The prompt carried the path, the format and the
read/write discipline. It looked something like "keep a small state file at
$HOME/workspace/.cron_state/blockers.json listing issue numbers already
reported, and skip those". Every run then depended on the model doing that
correctly. Two automations given the same path shared state without either
knowing. Nothing could inspect or reset what either believed.
Use the automation_state tool instead. Say what you want remembered in plain
language — "skip the issues you already reported". The agent stores it under
the automation it is running as. The scope is not a parameter, so one automation
cannot read or clear another's state. The tool is unavailable on an ordinary
chat turn, because there is no automation to scope it to.
State is bounded: 64 KB per automation, 16 KB per value, 512 keys. An oversized write is refused rather than truncated, so the agent learns and can adapt. You can read it and reset it from the automation's detail pane.
When to notify me
Whether a run's outcome reaches you used to be a request inside the prompt — "if nothing new, stay silent". That is a hope. The turn deciding whether to stay quiet is the same turn that wants to report.
| Setting | Behaviour |
|---|---|
Always | Every result is delivered. The default, and what automations did before. |
On change | Delivered only when the result differs from the last one delivered. |
On failure | Delivered only when the run fails. |
Never | Never delivered, not even on failure. The run history still records what happened. |
A failure is delivered under every setting except Never: a setting chosen to
reduce noise was not chosen to hide a failure. An empty result is delivered
under none of them, including Always. On change compares a whitespace-
normalised fingerprint, so a model that reflows the same answer does not count
as a change.
The comparison is kept where the model cannot reach it. An automation able to
edit the record of what it last said could talk itself past an On change
setting.
Which agent runs it
A job runs with the deployment's default agent unless it names another one. That
is what every job does until you declare agents in agents.named. When you
have, the editor grows a Runs as field listing them. A deployment that names
no agents does not show the field at all.
Naming an agent is not a hint the run may ignore. The agent is the ceiling:
- The turn sees only the tool groups that agent declared. A tool outside them is
not in the prompt. Calling it by name is refused rather than quietly allowed.
An agent scoped to
serverscannot draw a diagram, whatever the message asks for. - The agent's own instructions (its
addendum) are in front of the job's, and they add to the platform's rules rather than replacing them. - The job may only narrow what the agent has. Naming an agent and then asking for something it does not carry is refused when you save. The refusal message names the agent. Otherwise naming a small agent would be the way around its configuration.
Why bother: the apt-package-check-daily job wants one host and
execute_on_server. An agent scoped to exactly that is a much smaller blast
radius than the default agent running the same prompt at 03:00 with nobody
watching.
If you remove an agent from config while a job still names it, that job's next run stops with the reason. It does not fall back to the default agent, because the fallback would widen a job somebody narrowed on purpose.
Skills to load
With an agent named, this picker narrows that agent's skills instead of the whole catalogue. Picking none loads the agent's list. Picking a skill the agent does not have is refused. Without an agent it behaves as it always has.
Pick none and every skill is summarised in the prompt, as before. Pick some and
those load in full while the rest are left out, which keeps a focused automation
focused. Picking from the list is also usually how you find out a skill exists
at all. The alternative is writing a raw curl for something a skill already
does.
This is about what the model is told, not what it can reach. A skill is prompt content and grants no tool and no capability. So this is focus and discovery rather than a security boundary. Skills marked always-on stay loaded regardless, since that is a decision made for the whole workspace.
Delivery and Reliability
Automation delivery is workspace-local. Scheduled jobs and local trigger deliveries use the same workspace as the gateway.
Local trigger messages are written to a durable queue. If the gateway is not running yet, the message waits in that workspace. If the linked topic is already running a turn, the trigger waits until the session becomes idle instead of being injected into the active turn.
The local trigger queue is at-least-once, not exactly-once. If the gateway exits after claiming a delivery but before the linked turn completes, the next gateway start requeues that delivery. Use an idempotency key when firing over HTTP, and otherwise make repeated trigger messages safe.
What happens when a run fails
This is the question worth answering before trusting an automation overnight, and the two subsystems answer it differently on purpose.
Local trigger deliveries retry with exponential backoff and jitter, starting at 2 seconds and capping at 5 minutes, for up to 10 attempts. Backoff matters because the common failure is a downstream that is briefly unavailable. Without it, ten attempts burn in about five seconds. The automation then gives up long before the target recovers. Jitter is full jitter, so several gateways retrying against one host do not re-converge into the same burst.
After the tenth attempt the delivery is dead-lettered to
<workspace>/triggers/failed. It is not lost, and it is not retried again on
its own.
Scheduled automations do not retry unless the job asks for them to. The default is off, so nothing changes for a job created before this existed. When a retry policy is set, a failed run is retried with the same backoff, and:
- a skipped run does not consume an attempt — something declined to run it, and that is not the job failing.
- a one-shot job with attempts left is not switched off.
- a retry displaces the next scheduled slot rather than sitting beside it. A recurring job does not skip ahead as though the failure had not happened.
- each attempt is its own entry in the run history, so three failures read as three failures rather than one with a counter.
Replaying a failed delivery
A dead-lettered delivery describes exactly what should have happened, so it can be re-run once the cause is fixed:
# What is waiting
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
http://127.0.0.1:8765/api/webui/automations/failed
# Re-run one of them
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
http://127.0.0.1:8765/api/webui/automations/failed/tdl_abc123/replay
A replay is a new delivery that names the original in its replay_of field. It
therefore reads as a replay in the audit trail rather than as a fresh event. The
original stays in failed/ as the record of what went wrong. Attempts reset,
because a replay means someone believes the cause is fixed. A replay passes the
same capability gates as the original. It is a new execution with a known
provenance, not a recording being played back. Having run once before is not a
reason to skip an approval.
Replaying needs an operator token. A trigger key can fire, and cannot replay.
Rehearsing and granting over HTTP
The two commissioning actions are operator routes, so a trigger key does not reach them:
# Rehearse one automation now. Previews every gated action and runs none.
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
http://127.0.0.1:8765/api/webui/automations/4aa2442d/commission
# Write the standing grant its finding proposed.
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
http://127.0.0.1:8765/api/webui/automations/4aa2442d/grant
Both are plain GETs, like every other route on this surface. The gateway serves
its HTTP routes from the WebSocket handshake hook. That hook reads a request
line and no verb. So a -X POST closes the connection instead of reaching the
route.
The grant route sends no body on purpose. It writes only what the
automation's own commissioning finding proposed, because a grant the caller
could name would be a grant the caller chose. It answers 409 in three cases.
The automation has no refused finding. No grant could satisfy the finding — an
inventory write, an unbounded host set. The stored proposal is not a grant. The
response repeats what was written and states what it covers, and the gateway
needs a restart to apply it.
Audit records
Each local trigger delivery writes a record under <workspace>/triggers/runs.
Each scheduled run appends to that job's run history. The WebUI detail pane
shows that history with its status, duration, error and the reason it ran —
scheduled, manual or retry. Run one gateway consumer per workspace. The
local queue is not a distributed multi-consumer queue.
Common Patterns
For a nightly report, ask from the target topic:
Every night at 9pm, review today's workspace changes and summarize anything I should handle tomorrow.
For a CI follow-up, create a trigger once:
/trigger CI follow-up
Then have your CI or webhook adapter call:
nanoinfra trigger <trigger-id> "Build failed on main. Inspect the logs and suggest the next fix."
For a local report script:
generate-report | nanoinfra trigger <trigger-id>
A digest that only speaks when something changed
The instinct is to ask for silence in the prompt. Do the opposite: tell the automation to
always report the full current state, and set delivery to On change. The platform then
compares this run's answer to the last one it delivered and stays quiet when they match.
Check open issues on the repo and list every issue that is currently blocking, with its
number and one line on why. If none are blocking, say so.
That reads as though it will notify every night, and it will not. Asking the model to stay silent instead fails on the empty result: no policy delivers one. A run that correctly stays quiet and a run that broke then look identical from the outside.
On change compares whole answers. Remembered state tracks items. Use the policy when you
want to hear about the situation only when it moves. Use state when you want each new thing
mentioned once and never again, even as other things come and go:
Report only blockers you have not reported before. Use automation_state to remember which
issue numbers you have already mentioned.
Reach for state when the two disagree. An issue closing would change the whole answer and re-notify you about the ones that remain. Per-item memory avoids that.
A check that should be invisible until it breaks
Set delivery to On failure. A failed run still reaches you — see
When to notify me. So this is a check you genuinely never hear from while
it is healthy:
Verify last night's backup completed and the archive is readable. Fail loudly if it is not.
Pair it with a retry policy if the thing it checks can be briefly unavailable. A host that was rebooting then does not read as a broken backup.
A monitor that fires the automation itself
Issue the trigger a key once, then let the alerting system call it. Send an idempotency key — a monitor that retries a timed-out POST is the normal case:
curl -H "Authorization: Bearer $TRIGGER_KEY" \
-H "X-Nanoinfra-Trigger-Message: Disk on db-01 crossed 90%. Investigate and report." \
-H "X-Nanoinfra-Trigger-Idempotency-Key: disk-db01-2026-08-21" \
https://nanoinfra.internal/api/triggers/trg_8K4P2Q9X/fire
Derive the idempotency key from the alert, not from the clock, so the same alert retried later is still recognised as the same alert.
Catching up after an outage
When the thing an automation talks to was down long enough to exhaust the retries, the work is waiting rather than lost. Fix the cause, see what is queued, and replay it:
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
https://nanoinfra.internal/api/webui/automations/failed
Replaying runs the work now, under the current gates. Read the message first: an alert about a condition that has since resolved is often better dropped than replayed.
Referencing a server or a diagram
An automation that names a thing should not re-find it on every run. In a chat that search is cheap and self-correcting. At 03:00 it is neither, and a rename or a second host whose name is a closer match silently changes what the automation touches.
Reference it instead. The editor's message field takes the same @server: and
@diagram: mentions the chat composer does, and the automation stores the id:
Check @server:db-01 for failed units and report anything that is not running.
The name in the text is for you. The id is what the automation keeps. A server renamed afterwards still resolves and shows its new name.
If a reference stops resolving, the automation refuses to run. It does not fall back to searching by name, because that reintroduces exactly the ambiguity the reference removed, with nobody watching. The failure is terminal rather than retried. A deleted server fails identically on every attempt, and spending the retry budget re-learning that only delays the notification. The notification still arrives on an automation that has been quiet for weeks. See When to notify me for the reason.
The editor shows a reference that no longer resolves as broken, while you are editing. That is the point of capturing it there rather than typing it by hand.
Asking for the automation in chat captures the reference too. A message like "report the uptime of @server:db-01 every five minutes" mentions a resource. The automation is then created holding that server's id, with no extra step. Everything above then applies to it. The id is resolved again on each run. A deletion stops the automation instead of sending it looking for a name.
Troubleshooting
If an automation does not run, check that nanoinfra gateway is running, the
automation is enabled, and it was created from a linked topic.
If a local trigger waits forever, confirm the command uses the same workspace or config as the gateway.
If a trigger message appears twice after a restart, treat it as expected at-least-once delivery and make the external message idempotent.
If you need to edit, pause, resume, rename, delete, or inspect automations, use the WebUI Automations view.
If an automation is saved disabled right after you create it, a commissioning run found it would be refused on its schedule. The finding in its detail pane names what and why, and Grant it writes the grant that would permit it.
If the grant is right and the run is still refused, two things outrank it. A
matrix cell set to deny shadows every grant, and the refusal says which key to
change. And a session that was already denied once is latched. The gate is
not consulted at all until an operator clears it from the WebUI. A correct grant
changes nothing until then.
If an automation runs and its remote command refuses, read the capability gate
policy. A cron job and a local trigger both run in an unattended context, and the
shipped policy refuses each unattended remote command. A standing grant is the fix.
Heartbeat is the exception, because its turn carries the interactive context. See
capability-gates.md#which-turns-count-as-unattended.
Related Docs
webui.md#automationsfor the browser management viewchat-commands.md#local-triggersfor/triggercli-reference.md#local-triggersfornanoinfra triggerconfiguration.md#gateway-heartbeatfor heartbeat settingscapability-gates.md#choose-a-posturefor the policy that covers an automationguides/long-running-ai-agent.mdfor long-running agent work