Skip to main content

OpenAI-Compatible API

nanoinfra can expose a minimal OpenAI-compatible endpoint for local integrations:

nanoinfra plugins enable api
nanoinfra agent -m "Hello!"
nanoinfra serve

Run the CLI check first. If nanoinfra agent -m "Hello!" fails, fix provider or config setup before debugging the API server. By default, the API binds to 127.0.0.1:8900. You can change this in config.json.

For setup help, see quick-start.md, providers.md, and troubleshooting.md.

Two ways to serve it​

From the gateway (api.enabled), which is one process for both:

{ "api": { "enabled": true, "host": "127.0.0.1", "port": 8900, "apiKey": "${NANOINFRA_API_KEY}" } }

nanoinfra gateway then listens on its own port and on api.port, sharing one agent loop, one MCP host and one connector host. Default is false, so upgrading opens no port.

As its own process (nanoinfra serve), which is right for an API-only deployment and for a host that runs nothing else.

Prefer the first when you already run a gateway. A second process has to reassemble every piece of runtime the gateway already has, and a piece that goes missing there surfaces in production. Until 1.7.4, serve never drained the agent's outbound event bus. A turn emitting more than a thousand progress events then stalled mid-flight, and left its HTTP request unanswered.

Authentication​

Local-only 127.0.0.1 usage does not require an API key. If you bind the API server to all interfaces with api.host: "0.0.0.0" or "::", nanoinfra requires api.apiKey. Otherwise startup fails to avoid exposing an unauthenticated agent endpoint on the network.

{
"api": {
"host": "0.0.0.0",
"port": 8900,
"apiKey": "${NANOINFRA_API_KEY}"
}
}

When api.apiKey is set, send it as a Bearer token on API routes. The health endpoint remains unauthenticated so local probes and load balancers can still check process health.

curl http://127.0.0.1:8900/v1/models \
-H "Authorization: Bearer $NANOINFRA_API_KEY"

Behavior​

  • Session isolation: pass "session_id" in the request body to isolate conversations. Omit it for a shared default session (api:default).
  • The turn is what follows your last reply. A client that keeps its own transcript may resend it. nanoinfra then drops the messages up to and including the last assistant one, because this server kept its own record of them. It joins the user messages after it into this turn. nanoinfra joins a system message too, because it frames the turn. On its own a system message is not a turn, so a request with no user message is a 400.
  • Fixed model: omit model, or pass the same model shown by /v1/models.
  • Streaming: set stream=true to receive Server-Sent Events (text/event-stream) with OpenAI-compatible delta chunks, terminated by data: [DONE]. Omit stream or set stream=false for a single JSON response.
  • File uploads: supports images, PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx) via JSON base64 or multipart/form-data (max 10MB per file).
  • One log line per request — method, path, status, duration, and the reason a request was refused. Never the body and never the Authorization header: a request carries the conversation, so a log holding it would be a second copy of the transcript.
  • API requests run in the synthetic api channel, so the message tool does not automatically deliver to Telegram, Discord or any other channel. To proactively send to another chat, call message with an explicit channel and chat_id for an enabled channel.

Example tool call for cross-channel delivery from an API session:

{
"content": "Build finished successfully.",
"channel": "telegram",
"chat_id": "123456789"
}

If channel points to a channel that is not enabled in your config, nanoinfra will queue the outbound event but no platform delivery will occur.

Endpoints​

  • GET /health
  • GET /v1/models
  • POST /v1/chat/completions
  • POST /v1/responses

The Responses API​

POST /v1/responses is the same agent in the other wire shape. It is for a client that defaults to the Responses API, and that you cannot configure to speak Chat Completions. Same loop, same tools, same gate. Only the request and event shapes differ.

curl http://127.0.0.1:8900/v1/responses \
-H "Content-Type: application/json" \
-d '{"input": "hi", "session_id": "my-session"}'

input accepts a plain string, message items, or a message's content parts (input_text, and input_image with a base64 data URL). Several items follow the same rule as the chat route. What follows the last assistant item is the turn. The rest is your copy of a history this server also keeps. That is what makes the shape a real client sends work. The Codex CLI's first request carries its environment block and your prompt as two user items. That is one prompt in two pieces, not a transcript.

When the two copies disagree, ours wins: a client that pruned its transcript gets an answer informed by what the server kept. A follow-up can also name the conversation explicitly:

{"input": "and then?", "previous_response_id": "resp_..."}

previous_response_id resolves through a bounded index, so a very old id may answer 404 — name the conversation with session_id instead, which never expires. store: false keeps the id out of that index. Naming two different conversations in one request (a session_id and a previous_response_id belonging to another session) is a 400 rather than a guess.

Streaming follows the Responses event sequence: response.created, response.in_progress, response.output_item.added, response.content_part.added, then response.output_text.delta per token, and each level closed in turn before response.completed. Every event carries a sequence_number. There is no data: [DONE] sentinel — unlike the chat route, response.completed is this protocol's terminator, and a failed turn ends with response.failed rather than a truncated stream.

What it deliberately ignores​

nanoinfra accepts two request fields and ignores them. It logs each one once, so a client waiting for behaviour it asked for can find out why it never came:

  • tools (and tool_choice). nanoinfra runs its own tools. Returning function_call items for the caller to execute would make it a model proxy. It would also route around the capability gate, the confined executor and the audit log. So the endpoint lets a Responses client talk to the agent. It does not turn nanoinfra into a model backend for the client's own agent loop.
  • instructions. The system prompt belongs to the deployment. A caller able to replace it could ask for a different agent than the one the operator configured.

Input items that imply the caller ran a tool — function_call, function_call_output, reasoning — are refused with a 400 for the same reason.

curl​

curl http://127.0.0.1:8900/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "hi"}],
"session_id": "my-session"
}'

File Upload (JSON base64)​

Send images inline using the OpenAI multimodal content format:

curl http://127.0.0.1:8900/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": [
{"type": "text", "text": "Describe this image"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBOR..."}}
]}]
}'

File Upload (multipart/form-data)​

Upload any supported file type (images, PDF, Word, Excel, PPT) via multipart:

# Single file
curl http://127.0.0.1:8900/v1/chat/completions \
-F "message=Summarize this report" \
-F "files=@report.docx"

# Multiple files with session isolation
curl http://127.0.0.1:8900/v1/chat/completions \
-F "message=Compare these files" \
-F "files=@chart.png" \
-F "files=@data.xlsx" \
-F "session_id=my-session"

Supported file types:

  • Images: PNG, JPEG, GIF, WebP (sent to AI as base64 for vision analysis).
  • Documents: PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx) (text extracted and sent to AI).
  • Text: TXT, Markdown, CSV, JSON, etc. (read directly).

Python (requests)​

import requests

resp = requests.post(
"http://127.0.0.1:8900/v1/chat/completions",
json={
"messages": [{"role": "user", "content": "hi"}],
"session_id": "my-session", # optional: isolate conversation
},
timeout=120,
)
resp.raise_for_status()
print(resp.json()["choices"][0]["message"]["content"])

Python (openai)​

from openai import OpenAI

client = OpenAI(
base_url="http://127.0.0.1:8900/v1",
api_key="dummy",
)

resp = client.chat.completions.create(
model="MiniMax-M2.7",
messages=[{"role": "user", "content": "hi"}],
extra_body={"session_id": "my-session"}, # optional: isolate conversation
)
print(resp.choices[0].message.content)