OpenAI-Compatible API
nanoinfra can expose a minimal OpenAI-compatible endpoint for local integrations:
nanoinfra plugins enable api
nanoinfra agent -m "Hello!"
nanoinfra serve
Run the CLI check first. If nanoinfra agent -m "Hello!" fails, fix provider or config setup before debugging the API server. By default, the API binds to 127.0.0.1:8900. You can change this in config.json.
For setup help, see quick-start.md, providers.md, and troubleshooting.md.
Two ways to serve it
From the gateway (api.enabled), which is one process for both:
{ "api": { "enabled": true, "host": "127.0.0.1", "port": 8900, "apiKey": "${NANOINFRA_API_KEY}" } }
nanoinfra gateway then listens on its own port and on api.port, sharing one agent loop, one
MCP host and one connector host. Default is false, so upgrading opens no port.
As its own process (nanoinfra serve), which is right for an API-only deployment and for a
host that runs nothing else.
Prefer the first when you already run a gateway. A second process has to reassemble every piece
of runtime the gateway already has, and a piece that goes missing there surfaces in production.
Until 1.7.4, serve never drained the agent's outbound event bus. A turn emitting more than a
thousand progress events then stalled mid-flight, and left its HTTP request unanswered.
Authentication
Local-only 127.0.0.1 usage does not require an API key. If you bind the API
server to all interfaces with api.host: "0.0.0.0" or "::", nanoinfra requires
api.apiKey. Otherwise startup fails to avoid exposing an unauthenticated agent
endpoint on the network.
{
"api": {
"host": "0.0.0.0",
"port": 8900,
"apiKey": "${NANOINFRA_API_KEY}"
}
}
When api.apiKey is set, send it as a Bearer token on API routes. The health
endpoint remains unauthenticated so local probes and load balancers can still
check process health.
curl http://127.0.0.1:8900/v1/models \
-H "Authorization: Bearer $NANOINFRA_API_KEY"
Behavior
- Session isolation: pass
"session_id"in the request body to isolate conversations. Omit it for a shared default session (api:default). - The turn is what follows your last reply. A client that keeps its own transcript may resend it. nanoinfra then drops the messages up to and including the last
assistantone, because this server kept its own record of them. It joins theusermessages after it into this turn. nanoinfra joins asystemmessage too, because it frames the turn. On its own asystemmessage is not a turn, so a request with nousermessage is a400. - Fixed model: omit
model, or pass the same model shown by/v1/models. - Streaming: set
stream=trueto receive Server-Sent Events (text/event-stream) with OpenAI-compatible delta chunks, terminated bydata: [DONE]. Omitstreamor setstream=falsefor a single JSON response. - File uploads: supports images, PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx) via JSON base64 or
multipart/form-data(max 10MB per file). - One log line per request — method, path, status, duration, and the reason a request was refused. Never the body and never the
Authorizationheader: a request carries the conversation, so a log holding it would be a second copy of the transcript. - API requests run in the synthetic
apichannel, so themessagetool does not automatically deliver to Telegram, Discord or any other channel. To proactively send to another chat, callmessagewith an explicitchannelandchat_idfor an enabled channel.
Example tool call for cross-channel delivery from an API session:
{
"content": "Build finished successfully.",
"channel": "telegram",
"chat_id": "123456789"
}
If channel points to a channel that is not enabled in your config, nanoinfra will queue the outbound event but no platform delivery will occur.
Endpoints
GET /healthGET /v1/modelsPOST /v1/chat/completionsPOST /v1/responses
The Responses API
POST /v1/responses is the same agent in the other wire shape. It is for a client that defaults
to the Responses API, and that you cannot configure to speak Chat Completions. Same loop, same
tools, same gate. Only the request and event shapes differ.
curl http://127.0.0.1:8900/v1/responses \
-H "Content-Type: application/json" \
-d '{"input": "hi", "session_id": "my-session"}'
input accepts a plain string, message items, or a message's content parts (input_text, and
input_image with a base64 data URL). Several items follow the same rule as the chat route.
What follows the last assistant item is the turn. The rest is your copy of a history this
server also keeps. That is what makes the shape a real client sends work. The Codex CLI's first
request carries its environment block and your prompt as two user items. That is one prompt
in two pieces, not a transcript.
When the two copies disagree, ours wins: a client that pruned its transcript gets an answer informed by what the server kept. A follow-up can also name the conversation explicitly:
{"input": "and then?", "previous_response_id": "resp_..."}
previous_response_id resolves through a bounded index, so a very old id may answer 404 — name
the conversation with session_id instead, which never expires. store: false keeps the id out of
that index. Naming two different conversations in one request (a session_id and a
previous_response_id belonging to another session) is a 400 rather than a guess.
Streaming follows the Responses event sequence: response.created, response.in_progress,
response.output_item.added, response.content_part.added, then response.output_text.delta per
token, and each level closed in turn before response.completed. Every event carries a
sequence_number. There is no data: [DONE] sentinel — unlike the chat route,
response.completed is this protocol's terminator, and a failed turn ends with response.failed
rather than a truncated stream.
What it deliberately ignores
nanoinfra accepts two request fields and ignores them. It logs each one once, so a client waiting for behaviour it asked for can find out why it never came:
tools(andtool_choice). nanoinfra runs its own tools. Returningfunction_callitems for the caller to execute would make it a model proxy. It would also route around the capability gate, the confined executor and the audit log. So the endpoint lets a Responses client talk to the agent. It does not turn nanoinfra into a model backend for the client's own agent loop.instructions. The system prompt belongs to the deployment. A caller able to replace it could ask for a different agent than the one the operator configured.
Input items that imply the caller ran a tool — function_call, function_call_output,
reasoning — are refused with a 400 for the same reason.
curl
curl http://127.0.0.1:8900/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "hi"}],
"session_id": "my-session"
}'
File Upload (JSON base64)
Send images inline using the OpenAI multimodal content format:
curl http://127.0.0.1:8900/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": [
{"type": "text", "text": "Describe this image"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBOR..."}}
]}]
}'
File Upload (multipart/form-data)
Upload any supported file type (images, PDF, Word, Excel, PPT) via multipart:
# Single file
curl http://127.0.0.1:8900/v1/chat/completions \
-F "message=Summarize this report" \
-F "files=@report.docx"
# Multiple files with session isolation
curl http://127.0.0.1:8900/v1/chat/completions \
-F "message=Compare these files" \
-F "files=@chart.png" \
-F "files=@data.xlsx" \
-F "session_id=my-session"
Supported file types:
- Images: PNG, JPEG, GIF, WebP (sent to AI as base64 for vision analysis).
- Documents: PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx) (text extracted and sent to AI).
- Text: TXT, Markdown, CSV, JSON, etc. (read directly).
Python (requests)
import requests
resp = requests.post(
"http://127.0.0.1:8900/v1/chat/completions",
json={
"messages": [{"role": "user", "content": "hi"}],
"session_id": "my-session", # optional: isolate conversation
},
timeout=120,
)
resp.raise_for_status()
print(resp.json()["choices"][0]["message"]["content"])
Python (openai)
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8900/v1",
api_key="dummy",
)
resp = client.chat.completions.create(
model="MiniMax-M2.7",
messages=[{"role": "user", "content": "hi"}],
extra_body={"session_id": "my-session"}, # optional: isolate conversation
)
print(resp.choices[0].message.content)