HTTP API (lev serve)
lev serve exposes a REST + WebSocket API in front of the daemon, so anything that
speaks HTTP can drive Leviath, including The Lair, the browser console.
lev serve --port 3000 --token "$(openssl rand -hex 16)" --cors https://leviath.devEvery route on this page is also published as a machine-readable OpenAPI spec, kept in lockstep with the server by a test, so a client generator or an agent can consume the contract directly.
Security model
A token is required. The server refuses to start without
--token <t>(orLEVIATH_API_TOKEN). Every request must sendAuthorization: Bearer <t>; WebSocket clients pass it as?token=<t>because browsers can't set WS headers. On shared machines prefer the environment variable: a--tokenvalue is visible to other local users in the process table (ps).CORS is closed by default. Pass
--cors <origin>(e.g.https://leviath.dev) or--cors "*"to allow a browser to call it cross-origin. The server does not check theHostheader, so the bearer token is what stands between a DNS-rebinding page and your local API. Never embed the token in a page served from somewhere else, and avoid--cors "*"on a machine that browses.Binds to
127.0.0.1by default.--host 0.0.0.0exposes it on your network. Without--tls-cert, that puts the bearer token on the wire in cleartext for anyone on that network to read. If the address is publicly routable, that is the open internet. See reaching a Leviath on another machine.--tls-cert/--tls-keyserve HTTPS instead of HTTP. Off by default, bring your own certificate; Leviath never generates one.GET /needs no token. It returns a fixed "Leviath is running." page and nothing else: no version, no run counts, no endpoint list. It exists so a certificate can be accepted in a browser tab; see the section below.--allow-adminmounts the mutating admin routes.GET /api/configandGET /api/mcp/serversare always available. The writes are only mounted with--allow-admin, and the route is genuinely absent without it rather than gated by a check inside the handler. What you get back depends on whether the path exists at all for another method:Without --allow-adminResponse PUT /api/config405, because GET /api/configis mountedPUT /api/mime·DELETE /api/mime405, because GET /api/mimeis mountedPOST /api/mcp/servers405, because GET /api/mcp/serversis mountedDELETE /api/mcp/servers/{name}404, because nothing else is mounted on that path POST /api/update405, because GET /api/updateis mountedPOST /api/models/probe404, because nothing else is mounted on that path --workdir-rootconfines agent workdirs.--no-remote-yoloforbids"yolo": trueand"allow": [...]on spawn, which are one lever rather than two.--no-remote-seed-commandsruns every spawn as if it carried"no_seed_commands": true, so a blueprint'sseed = { command = ... }regions, which execute at spawn before any approval prompt, never run for a run that arrived over the API.lev runon the host is unaffected;[security] allow_seed_commands = falseis the machine-wide version.
Caution
lev serve runs LLM-driven tools with whatever permissions the blueprint grants. Treat it as
trusted-network only unless hardened. See Security.
Limits
The server holds a bounded number of requests in flight and gives each one a deadline. Both have
a default, both can be set in the config file or on the command line, and 0 switches either
off. A flag wins over the config file, and the config file over the default.
| Limit | Default | Flag | Config key | Over it |
|---|---|---|---|---|
| Requests in flight | 64 | --max-concurrent-requests <N> |
[serve] max_concurrent_requests |
503 at once, not queued |
| Seconds per request | 30 | --request-timeout-secs <SECS> |
[serve] request_timeout_secs |
408, and the handler is dropped |
| Bytes per request body | 32 MiB | none | [serve] max_upload_bytes |
413; this is what bounds a multipart upload |
Both answers carry the usual {"error": "..."} body. The websocket routes (/ws and
/ws/agents/{id}) are outside both: a subscription is meant to stay open. They are the only
place the API streams, so every other route has built its whole body before the deadline could
cut it. An unauthenticated request takes a slot while it is being refused, so a flood without a
token is refused at the cap like any other.
Three routes take a slot but have no deadline, because each waits on something slower than the default by design. Dropping one partway does damage a late answer would not:
| Route | What it waits on |
|---|---|
POST /api/mcp/servers/{name}/login |
The operator at the consent page, up to 300 s |
POST /api/mcp/servers/{name}/test |
The MCP server's handshake and tool listing, under the MCP client's own 30 s and 120 s deadlines. |
POST /api/doctor/live |
Two billed provider calls and a throwaway run, up to the doctor's own 90 s. |
Each is bounded by the deadline named in the table, so none can hold its slot forever. Dropping the login request closes the loopback listener the browser redirects to.
Neither limit is a ceiling on the runs behind the API. A spawn whose daemon takes a minute still
spawns; the route answers as soon as the daemon has accepted it. GET /api/config reports the
values in force under limits.max_concurrent_requests and limits.request_timeout_secs, so a
client sees what this server resolved rather than the default it would guess.
Reaching a Leviath on another machine
The short version: http:// only works on loopback. Everything else needs HTTPS or a tunnel.
A browser treats http://localhost and http://127.0.0.1 as potentially trustworthy, which is the
only reason the default setup works from a page served over HTTPS. Every other address is blocked,
and a LAN address is blocked exactly like a public one. http://192.168.1.50:3000 fails the same way
http://203.0.113.10:8080 does:
Mixed Content: The page at 'https://leviath.dev/lair' was loaded over HTTPS, but requested an
insecure resource 'http://203.0.113.10:8080/api/config'. This request has been blocked.Two things that are not the problem, because they are what people reach for first:
- It is not CORS. The request is killed inside the browser before it is sent, so it never reaches
Leviath and
--corsis never consulted. No response header on either side lifts a mixed-content block. - The site cannot fix it. leviath.dev is HTTPS-only, and an HTTPS page may not call
http://.
Pick whichever of these suits you.
mkcert, if the browser and Leviath are on machines you control
The best outcome: a certificate that is fully trusted, with no interstitial and nothing to accept. mkcert installs a local CA into your OS and browser trust stores and will issue for a bare IP.
TRUST_STORES=system,nss mkcert -install # once, on the machine running the BROWSER
TRUST_STORES=system,nss mkcert 192.168.1.50 # on the machine running Leviath
lev serve --host 0.0.0.0 --port 3000 \
--tls-cert ./192.168.1.50.pem --tls-key ./192.168.1.50-key.pem \
--cors https://leviath.dev --token "$LEVIATH_API_TOKEN"The TRUST_STORES=system,nss prefix limits mkcert to the OS and browser stores, which are the
only ones The Lair needs. Without it, mkcert 1.4.4 also looks at the Java trust store whenever
JAVA_HOME is set, on both commands. It then stops with failed to execute "keytool -list" and
Keystore file does not exist when that directory has a keytool but no lib/security/cacerts.
That is the layout brew --prefix openjdk has
(mkcert issue #472). If the install command
hit that error, the CA was already in your OS and browser stores before the Java step ran, so
rerun only the certificate command, with the prefix. Pointing JAVA_HOME at the real JDK home,
$(brew --prefix openjdk)/libexec/openjdk.jdk/Contents/Home, fixes it for good.
Installing a CA into your trust store is a real trust decision: anything holding that CA's key can
issue a certificate your browser will believe. mkcert keeps the key on the machine that made it.
Tailscale, for a publicly-trusted name
tailscale cert issues a real certificate for your *.ts.net hostname, so nothing needs installing
in a trust store and the port never faces the internet.
tailscale cert my-box.tail1234.ts.net
lev serve --host 0.0.0.0 --tls-cert my-box.tail1234.ts.net.crt \
--tls-key my-box.tail1234.ts.net.key --cors https://leviath.devSelf-signed, as a fallback
Works, with one manual step and one caveat.
openssl req -x509 -newkey rsa:2048 -nodes -days 365 \
-keyout key.pem -out cert.pem -subj "/CN=leviath" \
-addext "subjectAltName=IP:192.168.1.50"
lev serve --host 0.0.0.0 --tls-cert cert.pem --tls-key key.pem --cors https://leviath.devThen open https://192.168.1.50:3000/ in a browser tab and accept the warning. That is what the
unauthenticated GET / page is for. The Lair's requests are subresource fetch calls, which get
no interstitial to click through, so the exception has to be established in a tab first. Afterwards
The Lair works.
Chrome discards accepted exceptions when the browser restarts, so this comes back. Firefox keeps them. iOS Safari is unreliable about it.
SSH forward, if you would rather not deal with certificates
Nothing to install on either end, and it puts you back inside the loopback exemption.
ssh -N -L 3000:127.0.0.1:3000 you@that-machineThen point The Lair at http://127.0.0.1:3000. Leave Leviath on its default 127.0.0.1 bind for
this. --host 0.0.0.0 is not wanted and only widens the exposure.
Compressed responses
A client that says it takes gzip or br in Accept-Encoding gets any body over a kilobyte
compressed, with Vary: Accept-Encoding on the response. A run listing is JSON with a great deal of
repetition, so over a tunnel that is roughly a tenth of the bytes. Byte ranges, images and event
streams are left alone, as is anything already carrying a Content-Encoding.
Auth flow
sequenceDiagram
participant Client
participant Serve as lev serve
participant Daemon
Client->>Serve: request + Authorization: Bearer <token>
alt token missing / wrong
Serve-->>Client: 401 Unauthorized
else authorized
Serve->>Daemon: control-socket call
Daemon-->>Serve: result
Serve-->>Client: 200 JSON
end
Endpoints
Base path /api; all JSON unless noted.
Every route below can answer 401 when the bearer token is missing or wrong, so a client has to
handle that on all of them rather than on a few. The body is a line of plain text, not JSON.
GET /, the fixed status page, is the only route that does not check a token.
| Method · Path | Purpose |
|---|---|
GET /api/runs · DELETE /api/runs |
List runs: paginated, sortable, searchable · prune many at once. See below |
DELETE /api/runs/{id} |
Delete one finished run's record, and its sub-agent runs. Not the same as cancelling it. See below |
GET /api/agents · POST /api/agents |
List runs (deprecated, use /api/runs) · spawn an agent. Reads the persisted records, so finished runs stay listed |
GET /api/agents/{id} · DELETE … |
Get one · cancel. Cancelling stops the work and keeps the record; see deleting runs for removing it |
GET /api/agents/{id}/result · /context |
The run's answer and log tail · current context window |
GET /api/agents/{id}/logs?stage=&stream=&tail= |
A run's logs. stage, stream and tail pick which stage, which stream, and how much |
GET /api/agents/{id}/context/history |
How the context window changed over the run, paginated |
GET /api/agents/{id}/stages |
The per-stage ledger: tokens, dollars and regions per stage. See below |
GET /api/agents/{id}/files |
List a run's files, or read one with ?path=. offset pages a large one. See below |
GET /api/agents/{id}/files/raw?path= |
A workdir file's bytes under its own content type, for an <img> or a download. See below |
GET /api/agents/{id}/blobs · /blobs/{sha256} |
The stored parts a run holds, and one part's bytes. See below |
| The three byte routes above | Also take a short-lived ?exp=&sig= link in place of the bearer token, which is what a browser needs for an <img src>. Mint one with the GraphQL API |
GET /api/agents/{id}/artifacts/{name} |
The bytes of one file the run handed back, by the name its answer lists. See below |
GET /api/agents/tree · /{id}/tree-status · /{id}/children |
Sub-agent tree + token roll-ups |
POST /api/agents/{id}/pause · /resume |
Pause a run · resume it |
POST /api/agents/{id}/message |
Steer a running agent. Takes files too; see attaching files |
GET/POST /api/agents/{id}/interaction |
Read / answer a pending question. See below |
GET/POST/PUT/DELETE /api/blueprints[/{name}] · /validate |
Blueprint CRUD + validation, paginated and searchable with q |
GET /api/config · PUT /api/config (admin) · POST /api/config/validate |
Read redacted config · write or clear keys (null clears one) · validate a key |
GET /api/models?provider=&refresh= |
Enumerate models with their token limits and where they came from. See the model catalogue |
POST /api/models/probe (admin) |
Ask an OpenAI-compatible server what it serves, before writing a gateway for it. See below |
GET /api/providers?quota=&refresh= · POST …/{name}/login (admin) · /logout (admin) · /check (admin) |
The providers that sign in with a browser instead of taking a key, and the sign-in itself. See below |
GET /api/tools?agent= |
What an agent here can actually call. See below |
GET /api/mime · PUT /api/mime (admin) · DELETE /api/mime (admin) |
Read the effective mime registry, and write to it. The writes need admin. See below |
GET /api/scripts?agent=&include= · GET/PUT/DELETE /api/scripts/{kind}/{name} · POST /api/scripts/validate |
Read and write the machine's Rhai. Writes need admin. See below |
GET /api/mcp/servers · POST … · DELETE …/{name} · GET …/{name}/status · POST …/{name}/login · POST …/{name}/test |
List, add, remove, check, log in, test. All but the two GETs need admin |
GET /api/doctor · POST /api/doctor/live (admin) |
The checks lev doctor runs, as data. See below |
GET /api/yolo · GET /api/yolo/{name} · POST /api/yolo/test · PUT /api/yolo (admin) |
The profiles behind --yolo=<name>. See below |
GET /api/update |
Whether anything newer exists, how this copy was installed, and the command that upgrades it. See below |
POST /api/update (admin) · GET /api/update/jobs/{id} |
Carry that plan out, and read where it got to. See below |
GET /api/fs/dirs?path=&hidden= |
One directory level of subdirectory names, for a folder picker |
POST /api/fs/dirs |
Make one directory: {"path": "<absolute parent>", "name": "<one segment>"} → 201 {"path", "parent"} |
GET /ws · GET /ws/agents/{id} |
Live event stream (all agents / one run) |
GET /api/exports/{id} |
Collect a bulk export. Signed, five minutes, no bearer token. See below |
POST /graphql · GET /ws/graphql |
The same server answered field by field, and the same events over a subscription. See GraphQL |
Several routes in that table take more explaining than a cell allows.
GET /api/agents/{id}/stages reports what each stage spent per visit as well as in total, and
whether the stage ran at all. The blueprint detail route carries the manifest, the regions and the
fan-out limits.
A PUT /api/config that changes a provider key, a gateway, default_provider,
override_model or fallback_model applies to the next run
spawned, with no daemon restart. An MCP server added or removed through /api/mcp/servers reaches
the next run the same way.
GET /api/models answers from a catalogue the server keeps. An OpenAI-compatible gateway's
detected models are listed under the gateway's name. provider narrows the listing to one; see
two providers, one model id. refresh=1 asks the providers again;
see the model catalogue. POST /api/models/probe takes
{"base_url", "api_key"?, "headers"?} and answers {"models": [ids]}, or a 502 carrying the
server's own error text.
On GET /api/providers, quota=true adds what each subscription has used; see
subscription usage. refresh=1 reads the accounts again.
PUT /api/mime adds or updates a row in mime_types.toml. It takes
{"mime_type", "family"?, "text"?, "tokens"?, "extensions"?, "magic"?, "stand_in"?, "check"?}.
Only the fields you send are changed, and a bad type or rule is a 400.
DELETE /api/mime?mime_type= takes a row out, and answers 404 when there is no such row.
GET /api/scripts covers the agent's tools, hooks, validators and mime checks, and the global
model providers and mime checks. include=candidates also lists the files nothing declares yet.
GET /api/fs/dirs takes absolute paths only, fenced by --workdir-root, and hidden=true
includes dot-prefixed names. POST /api/fs/dirs uses the same fence, answers 409 when the
directory already exists, and is announced as fs.mkdir.
The doctor routes
GET /api/doctor is lev doctor --offline: config, search and resolve, nothing billed.
POST /api/doctor/live runs the whole chain, which is two billed calls and a throwaway run. It
answers 409 while one is already going. A failing check is ok: false inside a 200, never an HTTP
error.
Spawning under a yolo profile
POST /api/agents takes "yolo": true for a plain unattended run, and "yolo_profile": "<name>" to run under a named profile from yolo.toml instead.
A profile implies yolo, so the two need not both be sent. A name the file does not have fails
the spawn with a 400 that lists the profiles it does have. Under --no-remote-yolo a profile is
refused with yolo and allow: the operator's flag says nothing about which one.
The run listing and GET /api/agents/{id} carry the name back as yolo_profile beside
unattended, so a console can show that a run is unattended under something rather than
unattended outright.
Capturing one run's prompts
POST /api/agents takes "capture_model_input": true to write that run's exact requests into its
journal, once per provider attempt. It applies to the one run and leaves every other alone, so it
is the switch to reach for when the question is about a single run rather than a machine.
Read the warning before you send it. A captured request is the whole prompt, holding whatever the run's context held: file contents a tool read, command output, and anything somebody pasted. There is no size cap, so the journal grows by roughly the context size per attempt. See capturing what went to the model.
[observability] capture_model_input does the same for every run on the machine, and either one is
enough. Read the captured requests back on InferenceAttemptOutput.modelInput over
GraphQL.
Answering a question
GET /api/agents/{id}/interaction is the request the run is parked on: its id, kind, prompt
and options, plus tool_name and tool_arguments on a tool approval. POST the answer with
the request's id as request_id and one of three fields. value is free text or an edited
document. choice_index is a multiple choice, zero-based. approved carries an optional scope
(once, stage or session) for a tool approval or a confirm. It answers 202 once the daemon
has it and 404 when nothing with that id is open.
A deny may carry feedback, a string the model reads as part of the tool result for the refused
call, so its next turn is a redirect rather than a guess:
{"request_id": "coder-1788924523-abc123-approve-1", "approved": false,
"feedback": "use git log, not git show"}The model sees [denied] User declined tool call 'bash'. Feedback: use git log, not git show.
Without feedback the result is the plain [denied] User declined tool call 'bash'. it always was.
feedback beside approved: true is a 400, because there is nothing to redirect. The same text
is what the tool approval's fifth option, "Deny with feedback", collects in the dashboard, and
what lev respond <id> --deny --feedback "..." sends. Announced as interaction.feedback.
On /logs, stage takes a stage index or all, and defaults to the current stage. stream is
either output, the assistant's own text, or logs, which carries tool calls, token counts and
errors. tail is a byte budget for how much of the end you get back.
Note
A run object carries both updated_at and last_progress_at. The first advances on a 30-second
heartbeat and stays fresh on a run that has stopped; the second moves only when the run does. Age
a run against last_progress_at. pid is always 0 and means nothing: the daemon hosts every run
in one shared world, so there is no process per run. If you are tracking slots from outside, read
reconciling an external work queue first.
What a run flags about itself
A run object carries a flags object: post-hoc diagnostics that tell an empty or degraded run
from a healthy one without parsing its logs. Counters like modified_file_count, searches_run
beside searches_empty, and gates_forced say whether the run actually did anything, and
empty_output sums up the verdict. flags.broken_scripts names each Rhai script the run needed
but could not use, such as an output validator
that threw. It is the same list lev ps renders as (broken script), and the key is omitted
while the list is empty. The flags live in the run's meta.json, so they come back wherever a
run object does.
How long a run has taken
Three spans, and they answer different questions. A run paused overnight is hours old and spent almost none of them working. Report one where the reader wanted the other, and a healthy run looks stuck while a stuck one looks healthy.
| span | key | what it means |
|---|---|---|
| age | age_secs |
How long since the run was launched. Says nothing about whether it has done anything |
| working | working_secs |
How long it actually spent working. Call this the run's duration |
| last moved | last_progress_at |
When it last actually moved. A health signal, not a duration |
last_progress_at is how you tell a wedged run from a slow one.
age_secs and working_secs are computed server-side and appear on every run object this API
serves: GET /api/runs, GET /api/agents, GET /api/agents/{id} and
GET /api/agents/{id}/children, alongside the raw stamps they come from. ?fields= selects them
like any other key. The same two keys, on the same definitions, come back from lev ps --json.
The working clock stops for everything that is not the run's doing: paused, blocked on a person,
parked until the machine is fixed, finished. It keeps running while the run is inferring, calling
tools, or held for its own fan-out workers and sub-agents. Each stage in stages.json keeps one of
its own on the same rule.
To track a live run between the daemon's writes, read the clock itself rather than the computed
figure, which was true at server_time:
"active": { "banked_secs": 412, "since": 1787720786 }banked_secs is the time from spans that have already ended. since is when the span in progress
began, or null when the clock is stopped. The working total is banked_secs, plus now minus
since whenever since is set.
active is null on runs written before this existed, and working_secs then falls back to
updated_at minus started_at. A finished run has since: null, so its total never moves again.
Statuses
A run's status is one word, and it is the same word everywhere: on the run itself, in a tree node,
on GET /api/agents/{id}/result, and on the agent_status frames coming off the WebSocket.
| Status | Means |
|---|---|
starting |
Accepted and being set up. No inference has been issued yet |
running |
Working: inferring, calling tools, or moving between stages |
waiting_input |
Parked. wait_reason says on what, and only some of those want a person |
paused |
Paused, by somebody or by Leviath. Resumes on request |
complete |
Finished, with nothing further to accept |
complete_interactive |
Every required stage is done and the run still takes follow-up input |
error |
Stopped by a failure. The run's error carries what went wrong |
cancelled |
Stopped from outside. Nothing went wrong, somebody decided |
Leviath pauses a run itself when something outside the run has to change first, and wait_reason
is then needs_setup. A paused run comes back paused after a daemon restart.
The engine keeps its own vocabulary inside the daemon, where a run that is going is idle or
active and a parked one is waiting. Those words used to reach the socket untranslated, so a
client watching /ws was matching on three words no route ever sent, and a status frame quietly did
nothing for it. They are translated on the way out now.
Two older spellings are gone with them: GET /api/agents/{id}/result and the two tree routes
rendered the status for a human reader, which meant WaitingInput and CompleteInteractive where
every other route said waiting_input and complete_interactive.
events.run_status in the capabilities list is how you tell. A server without it sends the
engine's words on agent_status and the older spellings on those three routes.
The status= filter on GET /api/runs stays looser than this list on purpose: it also takes
waitinginput and Waiting-Input, so a status read off any response can be handed straight back as
a filter.
Region kinds
A region's kind, wherever one appears (GET /api/agents/{id}/context, its history, the blueprint
detail route), is the word the blueprint's own TOML uses: pinned, temporary, clearable,
sliding_window, compacting, compact_history, hashmap, checklist, custom.
Context snapshots written by an older daemon say sliding and history for the two multi-word
kinds, and those files stay on disk, so accept both spellings wherever you render one.
context.region_kinds says a server writes the blueprint's words.
Listing and searching runs
GET /api/runs returns a page, not the whole list:
{ "items": [{ "meta": { "run_id": "…" }, "highlights": [] }],
"next_cursor": "7b2276…", "total": 340, "server_time": 1785869070 }Pass next_cursor back as cursor and loop until it comes back null. Do not count pages against
total. It is what matched at the moment of that one request, and runs are being created and
finished underneath you.
Paging is keyset rather than offset, because an offset into a list that is changing skips and
repeats items, and does it most often at the head. sort=started_at is the default because it is
the only sort key that never changes. updated_at moves on the daemon's 30-second heartbeat, so
every live run shifts under a walk. A run whose sort value changes mid-walk can be missed or
repeated. To poll for what changed, use since= with no cursor rather than deep-paginating.
since= filters whichever field sort names, and is inclusive. Pass the previous response's
server_time and you may see one item twice, which is the safe direction when the granularity is
whole seconds.
Two parameters exist so a browser client does not have to make N requests: ids=a,b,c fetches exactly
those runs, and fields=run_id,status,title trims each one. Ids that no longer exist come back in
missing rather than failing the request.
Listing by place in the tree
A run's sub-agents are runs, so a listing that pages by runs is not paging by the rows a console
draws when it nests workers under the run that started them. At limit=50, seven visible rows and
forty-three workers hanging off them is a real page.
parent= fixes that from the server side:
| Value | Keeps |
|---|---|
| omitted | Every run, sub-agents included. What this route has always returned |
none |
Only runs nobody started. What a top-level list wants |
| a run id | That run's direct children, one level down |
total then counts what you asked for, which is what makes it worth printing beside a list: 382
under a sidebar drawing forty rows is comparing runs the reader can see against runs they cannot.
parent=<run_id> is also the paged, sorted, searchable form of
GET /api/agents/{id}/children, which answers the same question in one unbounded
array. A fan-out of two hundred workers has no windowed form there.
A run id that names nothing gives an empty page rather than a 404: a run with no children yet is
a normal answer, not a missing resource. none is the only keyword, and no run can collide with it,
since a run id is <agent>-<timestamp>-<hash>.
Announced as runs.parent. Without it, page until enough top-level rows exist to fill the viewport
and filter client-side, which is what The Lair does today.
Search
q= is a case-insensitive substring. It is not a regular expression, there are no boolean
operators or phrase quoting, and case folding is ASCII-only.
q_in= chooses where to look, defaulting to meta,files:
| Source | Looks at | Cost |
|---|---|---|
meta |
title, task, agent name, workdir, run id, error, metadata values | free |
files |
the paths the run recorded modifying | free |
context |
the run's current context window | one file read per run |
logs |
the tail of each stage's logs | two reads per stage per run |
journal |
the whole run journal: tool calls and context history | one file read per run |
The last three read from disk, which is why they are opt-in. Surface them as a "search inside
runs" toggle rather than making every keystroke pay for them. They also stop after a bounded number
of runs, newest first. When that happens the response says scan_truncated: true and sets total
to null, because a count taken from a partial scan would be read as fact.
Matching items carry highlights saying why they matched: the field, a snippet, and the stage
where there is one, which you can pass straight to /logs?stage=. This is the part that cannot be
done in the browser, because The Lair never holds a run's transcript.
One honest limit: the deep sources match the raw JSON on disk, so a query containing a quote, a backslash or a newline may not match text that does contain it.
descendant_of=<run_id> is the other tree question: every run under that one, at
any depth, and not the run itself. parent= is one level, which is what a tree
view expands a node with; descendant_of= is the flat read of a whole fan-out,
which is what a cost roll-up or an export wants. Setting both is a 400, because
they name two different sets and the one you meant is not recoverable from the
pair.
parent=sub is the mirror of parent=none: only runs somebody started, at any
depth. blueprint=<name> narrows to the runs of one blueprint, by the name each
run recorded, and composes with everything above.
Deleting runs
Cancelling and deleting are different verbs on purpose. DELETE /api/agents/{id} stops the run and
leaves everything it wrote; DELETE /api/runs/{id} removes the record. One stops the work, the
other forgets it happened.
Deletion is real and irreversible. The run's directory goes, transcript included. That is the point of the route: a "Delete" button that only hid the run in one browser would tell somebody clearing a sensitive transcript that it was gone when it was not.
DELETE /api/runs/deep-researcher-1786839472-d908ad2d9455
→ 204- 409 if the run is still going. Removing a directory out from under a running agent is a much larger feature than this, so cancel it first and delete it after.
- 404 if it is already gone, so a client that lost the response to its own delete can send it again instead of treating a missing run as a failure.
Sub-agent runs go with their parent
A fan-out worker and a sub_agent spawn are runs of their own, but they exist because something
started them, and they are drawn nested under it. Deleting the parent deletes them too.
Leaving them behind was not a matter of a few stale rows. A client that nests runs under their parent has nowhere to draw a run whose parent is missing, except the top level. Deleting a research run with nine workers under it emptied one row and promoted nine.
The walk only goes downwards. Deleting one worker out of a fan-out is an ordinary thing to do and leaves the run that started it, and the workers beside it, exactly where they were.
A live sub-agent is a 409 on the parent's delete, and the reason names the run to cancel. Half a tree is not a state anything downstream knows how to read.
A run whose meta.json will not parse is a 409 too, overridden with ?force=true:
DELETE /api/runs/{id}?force=trueA record that cannot be read says nothing about whether the run finished, and "cannot read it" must
not quietly read as "finished". That is exactly what a live run looks like to a binary whose
RunMeta has moved on. Such a run is also skipped by the listing, which would leave it both
invisible and permanent, so the escape hatch stays; it is something you type rather than
something that happens to you. The bulk route never forces.
Clearing out old runs
One request per run is its own problem once there are a few hundred, so there is a bulk form. It takes either an age or an explicit list:
DELETE /api/runs?before=1785869070
DELETE /api/runs?ids=run-a,run-b,run-cbefore is a unix timestamp and matches updated_at; only finished runs are considered. ids is
capped at max_ids from GET /api/config, the same cap as the batch fetch. Sending neither is a
400 rather than "every run". A bulk delete with no predicate is far more likely to be a client
that failed to build its query than somebody asking to erase the machine's history.
Partial success is the normal outcome, not an error. A sweep that runs into one live run has still correctly deleted the rest, so the response is a 200 with a verdict per run:
{ "deleted": ["run-a", "run-c"],
"skipped": [{ "id": "run-b", "reason": "Run 'run-b' is Running; cancel it before deleting it" }] }Read skipped when the list does not empty. A run that is still going and a run that was already
gone are both non-deletions, and only the reason tells them apart. Use the single-run route when you
want a status code per outcome instead.
deleted can hold ids you never named: every run named takes its sub-agent tree with it here too,
and those are runs that are now gone. Naming a parent and one of its own children in the same
request is fine. Each is deleted once, and the child is reported as deleted rather than skipped as
missing.
Exporting the whole store
Paging five thousand runs is a hundred requests, and a client that wants the lot wants it once. So there is an export. It is started over GraphQL, because that is where the filter already lives:
mutation {
startRunExport(request: {
filter: { status: { eq: COMPLETE } }
fields: ["run_id", "status", "cost_usd"]
}) {
export { id status }
}
}The answer comes back before the file exists. That is the point of a job: the request returns in milliseconds however large the store is. Poll it, and fetch it when it is ready:
{ runExport(id: "export-1789865498-0") { status written error downloadUrl } }status is QUEUED, RUNNING, COMPLETE or FAILED. written counts the runs on disk so far.
downloadUrl is null until the export is complete, and then it is a signed link to
GET /api/exports/{id}, good for five minutes. That link needs no bearer token, so it can be the
href of a download button.
The filter is the run listing's own, so a predicate that works for one works for the other.
fields narrows each row to the top-level run fields you name, using the same names the listing
uses. A name no run carries is a 400 rather than a column quietly missing from the file.
The body is JSONL: one JSON object per line, not one array. A reader can start on it before the writer has finished, and neither side ever holds the whole store in memory. Ten thousand runs is a loop over ten thousand lines at either end.
{"run_id":"run-a","status":"complete","cost_usd":0.0142}
{"run_id":"run-b","status":"complete","cost_usd":0.0071}A file is kept for one hour, then removed along with its job record. Neither outlives the other, so
an id that has expired and one that was never started both answer 404. While the export is still
running, or after it failed, the file route answers 409 and runExport carries the reason.
Where a run's cost went
GET /api/agents/{id}/stages returns one record per declared stage, in blueprint
order:
{
"run_id": "analyst-1786409275-d17e8f82",
"stages": [
{ "name": "plan", "status": "complete", "entered": true,
"prompt_tokens": 8420, "completion_tokens": 610,
"cached_tokens": 6100, "cache_write_tokens": 240,
"cost_usd": 0.0412, "unpriced_calls": 0, "cost_is_exact": false,
"cost_priced_usd": 0.0412,
"visit_count": 1,
"visits": [
{ "entered_at": 1786409280, "left_at": 1786409461,
"prompt_tokens": 8420, "completion_tokens": 610,
"cached_tokens": 6100, "cache_write_tokens": 240,
"cost_usd": 0.0412, "unpriced_calls": 0, "cost_is_exact": false,
"cost_priced_usd": 0.0412,
"active": { "banked_secs": 181, "since": null } }
],
"models": [
{ "provider": "anthropic", "model": "claude-opus-5" },
{ "provider": "openrouter", "model": "anthropic/claude-opus-5" }
],
"region_tokens": { "task": 24, "data_preview": 4004 },
"runaway_warned": false },
{ "name": "error_recovery", "status": "skipped", "entered": false },
{ "name": "answer", "status": "complete", "entered": true }
]
}Five things here are not derivable from any other route.
entered says whether the run was ever in that stage. The alternative is to
fetch context/history and diff consecutive snapshots to see which stages
produced entries. That is expensive, because every point carries a whole context
window. It is also wrong in the case that matters: a stage that ran and wrote
nothing to any region leaves no trace to find. status: "skipped" is the same fact
stated from the other side, and means the run finished without reaching this
stage, as distinct from "pending" on a run that is still going.
The per-stage cost split. The run-level totals are on the run record; which
stage spent them, and the cache read/write split within a stage, are only here.
A stage showing no cache reads cannot be told apart from one paying to write a
prefix nothing reuses without cache_write_tokens.
cost_usd is what that stage spent. It means exactly what it means on a run:
nullis unknown, never free. Some call in that stage was served by a model with no reported cost and no rates the daemon knows, so any total would understate by an unknown amount.unpriced_callssays how many.cost_is_exactsays which number you have.truemeans every priced call carried the provider's own figure: the invoice.falsemeans at least one was reconstructed from published rates, which is arithmetic on numbers that drift for reasons outside the daemon: negotiated pricing, a gateway's margin, a request rerouted to another backend.cost_priced_usdis the priced subtotal, kept even whilecost_usdisnullso a resumed run does not restart its accounting from zero. It is not a substitute forcost_usd: showing it while calls went unpriced is exactly the partial total that looks authoritative and is not.
Do not multiply the tokens by a rate card of your own. Pricing is the daemon's job, deliberately. A rate card in a console produces a fourth answer that disagrees with the run's figure, the stage's, and the provider's, and none of the four says which is wrong.
Every call a run bills is counted against the stage it was made in. That covers
the stage's own turns, the compaction calls that summarize its context when the
window fills, and the routing call it makes at its own boundary to choose where
to go next. The one exception is the run's title call, which happens once at
spawn beside the run rather than inside any stage of it. The stage costs can
therefore sum to slightly less than the run's own cost_usd.
visits splits a stage by each stay in it. The record above accumulates
across revisits. That is the right total for the stage and the wrong shape for a
graph of the path a run took, where a stage entered twice is two nodes. Each
entry covers one entry into the stage: entered_at, left_at (null on the
visit in progress), the same four token counts, the same four cost fields, and an
active working clock of its own. That clock follows the rule described under
how long a run has taken.
A stage that loops back to itself starts a new visit; iterations within one stay
do not. visit_count counts every entry, and the list stops at 128. So
visit_count > visits.length means the per-visit split is partial and the
accumulated figures on the record are the complete ones. visits is empty on a
stage the run never entered, and on records written by a daemon older than this
field, which is the other reason to keep falling back to the stage record itself.
models is what that stage actually ran on. The run's own model is the
entry stage's resolution and is never rewritten to follow the stages after it,
so it answers for one stage of the run and no other. This answers for each.
It is a list because a stage that fails over runs on more than one. The first entry is where the stage started, the last is where it ended up, and one entry means it never moved. A pair appears once however many calls it served, so read it as what ran rather than as how often.
A stage lists a pair once a call against it came back. So models is absent on
a stage the run never entered, on a stage whose first call is still in flight,
and on a stage whose only provider could not be reached at all. Choosing a model
is not running on one, and there is nothing here to read as a stage's intended
model.
region_tokens is what decides whether a region is earning its place. It is
the largest each region reached while that stage was active. This is the number to
look at before trimming a layout.
runaway_warned is set when a stage's per-call prompt passed four times its
first call, which is the shape of a region accumulating without a cap.
The list is bounded by the blueprint's stage count, so it is not paginated. A run that has not reached its first stage boundary returns an empty list rather than a 404. The run exists and has nothing to report yet.
lev stages <run-id> prints the same ledger as a table, --visits breaks each
stage into its stays, and --json is this shape read straight off disk.
Note
entered is false for every stage of a run recorded before Leviath tracked
it, because the field is not in those files at all. Read it together with
status: a stage recorded complete with tokens against its name ran,
whatever entered says on an old run. models is absent on every stage of
such a run for the same reason, and must not be filled in from the run's
model, which is the entry stage's and wrong for every stage after it.
Attaching files
A run takes files three ways, and all three end as typed parts on the region they
were aimed at, exactly as lev run --attach sends them.
multipart/form-data on POST /api/agents and POST /api/agents/{id}/message carries the bytes
themselves. A request field holds the JSON the route takes as a plain body. After it come any
number of file fields named part (bound for the task region, or the message's region) or
part:<region>, each with a filename and a Content-Type:
curl -X POST http://localhost:3000/api/agents -H "Authorization: Bearer $TOKEN" \
-F 'request={"blueprint":"storyteller","task":"a 30 second trailer"}' \
-F 'part:voice_samples=@voice.wav;type=audio/wav' \
-F 'part:storyboard=@frame1.png'POST /api/agents/{id}/interaction takes both forms too, for a text answer: the files land beside
the words in the tool result, and a choice or an approval with files is refused with 400.
A JSON body instead names files already inside the run's working directory under parts, each
{ path, region?, name?, mime_type?, deliver?, caption? }. And a @path token inside task, a
region's text, or a message names a workdir file the same way; the text keeps the token, so the
model reads the same name the part carries. A path that escapes the working directory is refused
with 403, a missing or empty file with 400, and a file over [serve] max_upload_bytes with 413.
The daemon types every part with its registry, so Content-Type and mime_type only need to be
right when the bytes and the name do not say. A part aimed at a region whose accepts excludes it
refuses the spawn with the region's list, and drops from a message with the text still delivered.
A run's parts
GET /api/agents/{id}/blobs lists every stored part the run's context holds: a user's attachment,
a file read_file stored, an image an MCP tool returned, an artifact the run submitted. Each item
carries sha256, mime_type, name, size, width, height, duration_ms, tokens, the
regions carrying it, and stored, which is false for a part the context names but the run
directory no longer holds. GET /api/agents/{id}/blobs/{sha256} serves one part's bytes under its
own Content-Type, so an <img src=...> pointed at it renders, and ?download=1 adds a
Content-Disposition: attachment carrying the part's name.
Both byte routes advertise Accept-Ranges: bytes and honour a single-range Range request, so a
player can scrub a video or a client can resume a download. Range: bytes=1024-2047 is answered
206 Partial Content with a Content-Range: bytes 1024-2047/<total> header and just those bytes;
an open end (bytes=1024-) or a suffix (bytes=-4096, the last 4 KiB) works too. A range that
starts past the end is 416 Range Not Satisfiable with Content-Range: bytes */<total>. A
malformed or multi-range header is ignored and the whole body served.
GET /api/agents/{id}/files/raw?path= does the same for any file inside the working directory,
typed by the registry from its bytes and name, where the JSON files route wraps text. A path the
directory does not hold but the answer lists as an artifact is served from the blob store instead,
so a file a model made and nothing wrote to disk answers here too.
GET /api/agents/{id}/artifacts/{name} is the route that follows an artifact the way the runtime
does: by the name the answer's artifacts list it under. It looks in the blob store by hash first
and the working directory second, and serves the bytes under the artifact's own mime type. It is
the one to use when a client has the answer in hand. The answer's artifacts carry each file's
path and, when the run could store it, its sha256, so the other two routes reach the same
bytes.
A run's files
GET /api/agents/{id}/files answers two different questions, and neither substitutes for the other.
source=modified (the default) is the run's own record of what it changed. It is free, but it is a
claim about the run rather than about the disk, and it is capped when recorded, so check
modified_files_truncated.
source=workdir reads the filesystem, one directory level per request; pass a directory as
path to descend. That bound is deliberate: a workdir containing node_modules cannot be
enumerated in one response, so walk it the way a file tree does.
Warning
modifying_tool_calls counts modifying tool calls, not files. A run that edits one file three
times records three. Do not subtract it from the entry count to get "how many more files";
that number is meaningless. Use modified_files_truncated, or source=workdir for ground truth.
Every listing entry carries name, path, is_dir, size, exists, outside_workdir, and
mime_type. The last is what the run's registry makes of the file from its name. A client can then
decide whether to render an image, or offer a file to a region that accepts a type, without a
request per row or a guess of its own. It is typed by extension only, not sniffed, and is empty for
a directory. GET .../files/raw types the same bytes, sniffing them, when an exact answer is
needed.
With ?path=<file> the response is the file's contents, unchanged from earlier versions. A listing
carries "kind": "listing", so check that field rather than guessing from the shape.
Reading a file larger than one response
One request returns at most 1 MiB. A run's dataset can be far larger than that, so read it a window
at a time with offset:
curl -H "Authorization: Bearer $TOKEN" \
"http://localhost:3000/api/agents/$RUN/files?path=data/dataset.csv&offset=0"Each response carries next_offset. Ask again from there until it comes back null, and
concatenate the windows to get the file back exactly.
An offset landing inside a multi-byte character is moved forward to the next boundary, and offset
in the response says where the window actually began. That is what keeps the pieces lining up. An
offset past the end of the file returns 416 rather than an empty window, so a loop cannot spin.
A whole-file read serializes exactly as it always has. offset is omitted when it is zero.
Fan-out limits
A fan-out stage has two caps, and neither is the stage's
max_iterations. max_workers is how many workers run at once and max_items is how many work
items the split may produce at all. GET /api/blueprints/{name} reports both for every fan-out
stage, resolved the way the daemon will apply them:
{
"name": "reviewer",
"fan_outs": [
{
"stage": "split_review",
"worker_stage": "review_worker",
"merge_stage": "deep_review",
"max_workers": 30,
"max_items": 30,
"on_worker_failure": "continue",
"results_region": "worker_findings"
}
],
"regions": ["…"],
"manifest": "…"
}max_workers is the default (30) when the manifest names none, and null when the stage is
unlimited. max_items is null when there is no ceiling. Whichever of worker_agent,
worker_stage or worker_query the stage uses is the one present. A blueprint that never fans out
has an empty list. blueprints.fan_outs in the capabilities list on GET /api/config says the
daemon reports this.
Changing a cap is a manifest write: PUT /api/blueprints/{name} with the manifest text, the stage's
max_workers or max_items set to the number you want, or to 0 for no cap at all. POST /api/blueprints/validate will tell you first if the value is not a whole number or is negative,
which are errors rather than quiet fallbacks. The workers still share the daemon's inference pool
([limits] max_concurrent_inferences, 8 by default), so an unlimited fan-out queues at the model
rather than running away.
Stage routing
GET /api/blueprints/{name} also carries stage_routing: one entry per stage that routes the
model's produced parts to a region by mime type (output_routing), or empties a region when it is
entered (context.reset). A console shows or checks them without parsing the manifest.
{
"stage_routing": [
{
"stage": "draw",
"output_routing": [{ "pattern": "image/*", "region": "artwork" }]
},
{
"stage": "describe",
"context_reset": ["conversation"]
}
]
}output_routing is ordered by pattern, and a part goes to the most specific match's region. Only
stages that do one or the other appear; a stage that does neither is left out, and a blueprint that
does neither has an empty list. Changing either is a manifest write through PUT /api/blueprints/{name}. blueprints.stage_routing in the capabilities list on GET /api/config
says the daemon reports this.
dependencies on the same detail route lists what the agent declares it needs before it runs (see
dependencies). Each entry has a name, a kind (mcp_server,
env, binary or script), whether it is required, and an optional remedy and description.
It also carries the kind-specific field (server and env, var, command, or check), and
installable for whether the blueprint says how to set it up. A console can show these and warn
before a spawn that would fail the dependency gate. blueprints.dependencies in the capabilities list says the daemon
reports this.
Yolo profiles
The profiles a run can be launched under with --yolo=<name> live in
yolo.toml beside the config, and these four routes are lev yolo over HTTP. Every one of them reads the file as it stands at that moment, the same way a
spawn does, so what they report is what the next run gets.
GET /api/yolo lists what is there:
{
"path": "/home/you/.leviath/yolo.toml",
"exists": true,
"profiles": [
{
"name": "careful",
"default": "ask",
"questions": "ask",
"checkpoints": "ask",
"gate": "auto",
"tool_rules": [1, 2, 0],
"shell_rules": [3, 1, 1]
}
]
}tool_rules and shell_rules are [allow, ask, deny] counts. A file that does not load comes
back with exists: true, an error naming the line, and no profiles, because that is what a
spawn naming one would be refused with. A missing file is exists: false with no error.
GET /api/yolo/{name} is one profile in full: {"name", "spec", "holds"}. spec is the profile
as parsed, with the same keys the file has. holds is the list of things this profile still puts
to a person, as lev run prints before a run starts. 404 for a name the file does
not have, and for no file at all; 422 when the file does not load.
POST /api/yolo/test asks what a profile would decide for one call, without running anything:
{
"profile": "careful",
"tool": "shell",
"command": "rm -r target/debug",
"workdir": "/home/you/project"
}command is for the shell; any other tool takes its arguments as an object. workdir is where
relative paths in the command resolve, defaulting to the server's own. configured (allow,
ask or deny) stands in for what the config layers resolve the tool to. Without it, that comes
from the config in force. kind (builtin, subagent, script, mcp) says where the tool
comes from for @group rules. Without it, the kind is guessed from the name. allowed: true
decides as if --allow <tool> had been passed. The answer:
{
"profile": "careful",
"tool": "shell",
"configured": "ask",
"policy": "allow",
"reason": "shell allow rule \"rm -r*\""
}policy is what the run would do: allow runs it without a prompt, ask opens the ordinary
approval prompt, deny refuses it. reason names the rule, the config, or the profile's default
that decided it. This is the same code path the daemon runs on a real call, so a decision here is
the decision a run would make. 400 for a kind or configured word that is not one of the
listed values, or arguments that are not an object.
PUT /api/yolo replaces the whole file: {"text": "<the file, as TOML>"}. The text is parsed
first, and a save that would not load is refused with 400 and the same message a spawn would
give, leaving the file on disk as it was. It answers with the listing GET returns. Admin only,
for the reason PUT /api/config is: a profile is a grant of permissions.
Asking how to upgrade
GET /api/update answers how this copy of Leviath was installed and what command brings it
up to date. Like GET /api/tools, it exists because the answer is a fact about the machine
that no client can work out for itself, and guessing it wrong is worse than not saying.
The body is exactly what lev update --check --json prints, from the same planner:
{
"version": "0.4.0",
"install_method": "scoop",
"channel": "stable",
"binary": {
"action": "run",
"commands": [["scoop", "update"], ["scoop", "update", "leviath"]],
"command": ["scoop", "update", "leviath"]
},
"agents": [],
"migrations": [],
"config_error": null,
"latest": "0.4.2",
"update_available": true,
"checked_at": 1787438706
}install_method is one of homebrew, scoop, cargo, script or unknown. binary.action
is either run, carrying a commands list of argv lists to run in order, or advise, carrying
a message to show instead. A cargo install copy is always advise: rebuilding it is a full
compile, which is not something to start on someone's behalf. So is a binary sitting somewhere
no installer puts one, where the honest answer is to point at the install docs.
Render binary.commands rather than composing your own. That is the whole point of the route:
a client that hard-codes one package manager's command is right for the users who happen to
share its author's machine and wrong for everyone else. Where a daemon does not announce
update.plan, send people to the install page rather than picking a package manager for them.
latest is the newest version on this copy's own channel, and update_available says whether
that is newer than the version it is running. checked_at is when the daemon last found out, in
unix seconds, so you can say how fresh the answer is rather than presenting an hour-old one as
current. All three are null together when the check has not run yet, could not reach the
network, or had no channel to ask about. That is one state, "cannot tell", which is the honest
thing to render. Treat a missing key as an older daemon and a null key as an answer.
The daemon looks this up on its own schedule and the route reports whatever the last lookup
found, so asking on every page load costs nothing and never waits. The lookup runs the same code
lev update does, against the releases published for each channel, so the console and the
terminal cannot come to different conclusions about the same binary.
Do the comparison with update_available rather than against version yourself. A client that
compares against a number it was built with only knows the stable line, so it reports a daemon on
alpha or beta as out of date for running something newer.
The route is read-only and available without --allow-admin. It works out what an update would
do and does none of it, makes no network call on the request path, and cannot run a command even
if asked.
Pressing the button
POST /api/update carries out the plan the GET prints: it runs binary.commands in order,
installs the blueprints the plan marks preselected, and applies the migrations. It needs
lev serve --allow-admin, which is the line it crosses and the read half does not. It runs a
package manager, replaces the blueprints in your agents directory and rewrites your config.
The body names which parts to do. Every field defaults to true, so an empty body is the whole
plan and a body naming one part leaves the others on. A field this route does not know is a 400
rather than a silent default:
{ "binary": true, "agents": true, "keys": true, "migrations": false }It answers 202 straight away, with the id to watch:
{
"job_id": "update-1787438706-1",
"status": "running",
"applying": { "binary": true, "agents": true, "keys": true, "migrations": false }
}An upgrade is a download and an install: a minute on a good day, and it can fail halfway. The
request does not stay open for it. Watch /ws, where each step change arrives as it happens:
{ "type": "update_progress", "job_id": "update-1787438706-1", "step": "binary",
"status": "running", "detail": "running `scoop update && scoop update leviath`" }step is binary, agents, keys or migrations, always in that order, and status is one
of running, done, skipped, advised or failed. The keys step respells blueprint keys
that changed name, in the blueprints you wrote rather than the bundled ones. The last frame is update_finished, carrying
the whole record so a client that connected mid-run needs no follow-up request. Both frames are
about the machine rather than a run, so /ws receives them and a per-run subscription does not.
GET /api/update/jobs/{id} answers that same record, for a client that would rather poll than
hold a socket open:
{
"id": "update-1787438706-1",
"status": "complete",
"steps": [
{ "step": "binary", "status": "done", "detail": "ran `scoop update && scoop update leviath`" },
{ "step": "agents", "status": "done", "detail": "installed researcher, coder" },
{ "step": "keys", "status": "skipped", "detail": "every blueprint uses the current key names" },
{ "step": "migrations", "status": "skipped", "detail": "not asked for" }
],
"restart_required": true,
"restart_hint": "the new binary is on disk, but this server and the daemon it talks to are still running the old one...",
"started_at": 1787438706,
"finished_at": 1787438771
}The last few runs are kept, so reading back after the fact finds the job rather than a 404.
One update runs at a time: a second POST while one is going is a 409 naming the job already
running, not a second package manager over the same binary.
What it will not do
binary.action == "advise" stays advice. A cargo install copy is a full rebuild of the
workspace, and a binary somewhere no installer writes is not something to guess at. Both are
yours to do, so the step is recorded as advised with the plan's own sentence and no compile is
started. That is neither a success nor a failure: the job carries on to the other two steps and
still finishes complete.
A blueprint you edited locally is never installed. Installing removes the destination directory
first, so it would take your edits and any file you added with them. lev update asks about each
one on its own and no flag covers it, and there is nobody to ask over HTTP. The agents step says
how many it left alone and why.
A binary step that fails stops the two after it, the same way lev update stops there: the
blueprints and the config worth having are the ones the new binary ships. A failed blueprint
install does not. It is named in the step's detail and the run carries on. Most of the blueprints
plus a named failure is a better place to be left than a step that gave up in the middle.
The restart
Upgrading replaces the binary on disk. The daemon answering the request is the old one and stays
the old one, and so does lev serve, until each restarts. A console that updates and then
reports the version it can see has told the truth in the least useful way possible.
restart_required is true when the binary step actually ran and succeeded, and restart_hint
carries the sentence to show. Say it; do not report the running version as the result of the
update. Restarting lev serve picks up the new binary, and lev daemon restart does the same for
the daemon. Any lev command also does that on its own, since the daemon's build marker is
checked before a run is spawned.
Tools and scripts
GET /api/tools answers what an agent on this machine can call, which is not a question a
client can answer for itself. Every entry carries a source:
source |
Means |
|---|---|
builtin |
Compiled into this Leviath. Every agent has it |
subagent |
A sub-agent tool, for an agent that may spawn children |
agent |
A .rhai in that agent's own tools/. Only that agent has it |
global |
A .rhai in ~/.leviath/tools/. Every agent on the machine has it |
Pass ?agent=<name> to include the fourth. Every entry also carries the description and the
arguments schema the model is given, so a picker shows what a tool takes without compiling
anything. Script-backed entries carry the path they came from and the capabilities they
requires. A separate skipped list carries the .rhai files that were found and cannot be offered,
with the reason each was passed over. A file with a syntax error is then told apart from a file
nobody wrote. MCP tools are not here: they depend on a server being reachable rather than on
anything installed, and /api/mcp/servers/{name} already answers for them.
GET /api/scripts is the same ground from the editor's side, over the six kinds of Rhai a machine
can carry: tool, region_hook, stage_hook, output_validator, mime_check and provider.
Only tools have a directory an agent owns: <agent>/tools/, plus the global one. The hooks and
the validator are named by path in the manifest and resolved against the agent's own directory.
The listing derives them from what the manifest declares, and the read and write routes address
them at <agent>/<name>.rhai.
A mime check is named by a mime row's check, and rows live in two
places, so the kind is listed from both. The operator's rows (mime_types.toml and
[mime_types] in the config) put their checks in the global half, resolved against the config's
directory. A blueprint's own [mime_types] puts its checks beside the agent's hooks. Address
one with ?agent=<name> for the blueprint's, or without for the operator's. Check
scripts.mime_checks in the capabilities list before offering the kind.
Every entry carries a declared flag, and an agent-scoped one also carries relative_path: where
the file sits relative to the agent's own directory, validators/a2ui.rhai, which is the spelling
that goes into a manifest. A global mime check carries it too, relative to the config's directory,
since that is the spelling that goes into the row. A global tool or a provider has no
relative_path, since nothing names either by path.
GET/PUT/DELETE /api/scripts/{kind}/{name} reads and writes one file, scoped by ?agent=<name> or,
with no agent, the machine's own directory for that kind. {name} is the file without its .rhai
extension, and it may be a relative path when the manifest declared one: percent-encode the
separator, so validators/a2ui.rhai is output_validator/validators%2Fa2ui. Every part of it may
hold only letters, digits, ., _ and -. The result has to land inside the directory the route
is fenced to, once symlinks are followed. A declaration that climbs out of the agent's directory
is left out of the listing, as is one that names something that is not a .rhai file. The
alternative would be reporting it under a name that would fetch a different file.
POST /api/scripts/validate takes kind and content and compiles without writing, so an editor
can check before saving instead of saving and waiting for a run to fail.
Offering a file nobody has named yet
The listing above answers "what will load", which is circular for a picker: a hook or a validator
appears once the manifest declares it, and declaring it is the thing the picker exists to do. So
GET /api/scripts?agent=<name>&include=candidates adds the other half, the .rhai files under that
agent's directory that nothing declares:
{
"kind": "unknown",
"name": "validators/draft",
"source": "agent",
"agent": "picker",
"path": "/home/you/.leviath/agents/picker/validators/draft.rhai",
"relative_path": "validators/draft.rhai",
"declared": false
}kind is unknown because nothing about the file says which of the four agent-owned kinds it is;
the declaration says that, and it has not happened yet. unknown is not a {kind} the read and
write routes accept, so a client picks a real one to open the file with. compiles is absent
for the same reason: which compiler would have to accept it is not yet decided. Write
relative_path into validator = "..." or a [stages.<name>.hooks] entry and the next listing
reports the same file as declared.
Without the parameter the listing is exactly what it was, so a client that reads it as "what will
load" keeps getting that. include takes a comma-separated list and refuses a token it does not
serve, rather than answering a misspelling with a short list. With no ?agent= it changes nothing:
both global directories are already listed file by file, so nothing there is undeclared. Check
scripts.candidates in the capabilities list before offering the picker.
The scan is bounded and stays inside the agent's directory: four levels deep, 128 directories and
256 files per request, .rhai files only, and no symlink is followed out of the agent's own
directory. A file whose name could not be written into a manifest, because of a space or a character
that is not a safe path component, is left out rather than guessed at.
Warning
PUT and DELETE are not mounted at all without lev serve --allow-admin, exactly like the
MCP add/remove routes and PUT /api/config. A .rhai file is executable code every agent then
runs, so a session that can write one can run code on the host. The GET routes stay open, so an
editor degrades to read-only rather than disappearing.
A write that does not compile is still saved, with compiles: false and the compiler's complaint in
the response. A draft is worth keeping, and a tool that does not compile is skipped at spawn rather
than breaking the agent.
Providers
A provider is the one kind no agent owns. It lives in ~/.leviath/providers/
and a stage reaches it by name, so these routes take no ?agent= and refuse one rather than
writing a file into an agent's directory that nothing would ever load. Providers are listed with or
without an agent, since the answer is the same either way.
Each listed provider carries a provider object with what its leading // @ comments declare:
description, default_model, max_context_tokens, max_output_tokens and supports_streaming.
The model catalogue
GET /api/models answers from a listing the server keeps, not from the providers on each
request. The providers are asked once per config, side by side, each given five seconds. A
complete listing is served for fifteen minutes, and one missing a provider's answer for one
minute. After that, the next request gets the list in hand and starts a refresh behind it.
Only the very first request for a config waits for the providers, and that wait is bounded.
A PUT /api/config starts the refresh for the new config as it returns.
Two headers say what you got. X-Leviath-Catalog-Age is how many seconds ago the listing
was built. X-Leviath-Catalog-Complete is true when every provider answered when it was, and
false when one timed out, errored, or could not be built. That provider's models are then
absent. ?refresh=1 asks the providers again and waits for them, for a settings page that
has just changed a key and wants to show the result rather than the memory of the old one.
Announced as the models.cached capability.
Two providers, one model id
GET /api/models returns a flat list, and each entry carries the provider
that serves it. That matters more than it looks: openai and codex serve
the same model ids. Both answer to gpt-5.5. They bill to entirely
different places: one to an API balance, one to a ChatGPT subscription.
So id is not a key. provider + id is:
[
{ "id": "gpt-5.5", "provider": "openai", "max_context_tokens": 400000 },
{ "id": "gpt-5.5", "provider": "codex", "max_context_tokens": 400000 }
]Measured against a live daemon signed in to both, that is four ids each served
twice: gpt-5.5, gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna.
A client keying a picker on id alone silently collapses those two into one
row, and whichever it kept decides what the user pays. Build the key the way a
blueprint names a model: provider/id, the same codex/gpt-5.5 that goes in
models = [...].
?provider= asks the server instead:
GET /api/models?provider=codex only the subscription's models
GET /api/models?provider=openai only the keyed onesA provider this machine has not configured lists nothing rather than 404ing. The set of providers is whatever the config has, so "no models" is the honest answer to asking about one it does not have.
pricing is null on a Codex entry, and on any other model whose provider
did not quote a rate in its listing. It reports what the listing said, not
what a run would be billed at, so a console should not read null as free.
For Codex it happens to be free, because a subscription has no per-call price
and a run on it reports cost_usd: 0. The two facts arrive by different routes.
See where a run's cost went for the figure that is
actually charged.
Signing in to a subscription provider
Most providers take an API key, which is a string a form can post. One does not: Codex bills a
ChatGPT subscription and its credential is an OAuth grant, taken in a browser and stored outside
config.toml entirely. PUT /api/config can turn it on and can never sign anybody in, so a
console that only wrote the flag would leave the user enabled and unable to run anything.
These four routes are the rest of it.
GET /api/providers what exists, and what state it is in
POST /api/providers/{name}/login start the browser flow (admin)
POST /api/providers/{name}/logout forget the grant (admin)
POST /api/providers/{name}/check prove the grant still works (admin)GET /api/providers is open to any caller holding the bearer token. It reports the account
address and the plan tier, the same facts lev auth status prints, and never a token. Two
providers sign in this way: codex (a ChatGPT plan) and grok (a SuperGrok or X Premium+ plan).
{
"providers": [
{
"id": "codex",
"display": "OpenAI Codex",
"enabled": true,
"signed_in": true,
"account": "someone@example.com",
"plan": "plus",
"expires_at": 1735689600
}
]
}Subscription usage
GET /api/providers?quota=true adds a quota object to each signed-in provider that is enabled:
{"report": {...}} with what the subscription has used, or {"error": "..."} when the account
could not be read. It is off by default because it is a reading of the accounts rather than of
this machine, and a console polls this route while a sign-in is waiting.
"quota": {
"report": {
"plan": "plus",
"windows": [
{ "label": "5h", "used_percent": 42.0, "resets_at": 1735700000 },
{ "label": "week", "used_percent": 12.5, "resets_at": 1736200000 }
],
"limit_reached": false
}
}A window carries used_percent when the provider reports a share (Codex), or used, limit
and unit when it reports amounts (Grok, in credits). balance is a prepaid balance as the
provider words it, when there is one. lev providers quota --json prints the same reports.
The accounts are asked side by side, each given five seconds, so the reading is bounded by five
seconds however many subscriptions are signed in. One that does not answer inside that carries
{"error": "the account did not answer within 5s"} and the rest of the reading still arrives;
nothing holds the response.
The answer comes from a reading the server keeps, not from the accounts on every request. A
complete reading is served for one minute and one missing an account's answer for fifteen
seconds, after which the next request gets the reading in hand and starts a new one behind it.
Two headers say what you got: X-Leviath-Quota-Age is how many seconds ago the accounts were
read, and X-Leviath-Quota-Complete is true when every one of them answered. complete names
no provider, so the per-provider error is the one to read; the header is there to say whether
asking again is likely to help. ?refresh=1 reads the accounts again and waits, for a "check
again" button, and means nothing without quota=true. Signing in or out gets its own reading
rather than the one taken before it, so a console does not have to invalidate anything itself.
Announced as the providers.quota capability. Safe to ask for whenever the providers page
opens.
enabled and signed_in are separate on purpose. Either can be true alone, they are set by
different routes, and the combination that breaks runs, enabled but not signed in, is the one a
console has to be able to see. expires_at is when the access token lapses, and it is refreshed
automatically well before that. It is here so a UI can show that the session is live, rather
than implying anybody has to act on it.
The flow
POST .../login does not hold the request open for the whole sign-in. It answers as soon as
there is a URL, which is immediately, and the flow carries on behind it:
sequenceDiagram
participant UI as Console
participant Serve as lev serve
participant Browser as Browser on the serving host
participant OpenAI
UI->>Serve: POST /api/providers/codex/login
Serve->>Browser: open the authorize URL
Serve-->>UI: 202 {status: waiting, authorize_url}
Note over UI: render the URL, in case no browser opened
Browser->>OpenAI: sign in
OpenAI->>Serve: redirect to localhost:1455/auth/callback
Serve->>OpenAI: exchange the code
loop until it settles
UI->>Serve: GET /api/providers
Serve-->>UI: signin.state = waiting
end
UI->>Serve: GET /api/providers
Serve-->>UI: signed_in: true, no signin state
The MCP login route does hold its request open, for up to five minutes. That is fine for a CLI and wrong for a browser. The tab has nothing to draw while it waits, no way to show the URL, and no way to give up without losing the flow. One extra poll buys a UI that can render the whole thing.
While it runs, the provider carries a signin object:
{ "state": "waiting", "authorize_url": "https://auth.openai.com/oauth/authorize?…",
"started_at": 1735689000 }and if it does not finish:
{ "state": "failed", "message": "could not listen on port 1455 or 1457 (…)", "at": 1735689100 }A success leaves no signin at all. The grant store is the answer, and a second copy of "signed
in" is how two answers come to disagree. A failure stays until the next attempt, so a console
that was not watching still finds out what happened.
Starting a second sign-in while one is waiting answers 409, with the first one's
authorize_url in the body. A console that lost the response can pick the flow back up instead
of having to cancel it.
The browser has to be on the serving host
The redirect goes to localhost:1455 on the machine running lev serve and nowhere else,
because that is the address OpenAI registered the public client id against. Leviath cannot choose
a different one.
For a console driving a daemon on the same machine, this is invisible: the browser opens and the
callback lands. For a remote daemon, the flow still starts and authorize_url still comes back,
but somebody has to open that URL on the daemon's host. Show it rather than relying on the
opener, which silently does nothing over SSH and in a bare console.
Checking, and why the check is real
POST .../check asks the account, not a table:
{ "status": "ok", "provider": "codex", "models": ["gpt-5.6-sol", "gpt-5.6-terra", …] }Codex's model list is compiled into the build, so a check that read it would answer "5 models"
for a revoked session, an expired one, and for never having signed in. This one makes an
authenticated call, so a 200 means the subscription really did agree, and the models are the
ones this plan can reach. A refusal is a 502 carrying the provider's own reason.
It also refreshes a lapsed access token on the way in, which makes it the cheapest way to keep a rarely-used sign-in alive.
Turning it on
Signing in and enabling are two separate acts, and logout deliberately does not do both.
Signing out is not the same as turning the provider off. PUT /api/config carries the setting:
{ "codex_enabled": true, "codex_reasoning_effort": "medium", "codex_verbosity": "medium" }codex_reasoning_effort takes none, minimal, low, medium, high or xhigh;
codex_verbosity takes low, medium or high. Both are validated before anything is written,
because the provider silently ignores a value it does not recognise. A typo would otherwise be
saved, read back by GET /api/config, and quietly do nothing.
ollama_enabled turns Ollama on, and GET /api/config reports it. It is
opt-in like everything else: off, no run registers it. Setting
ollama_base_url counts as choosing it too, and the GET reports true for
either. A console asking "is Ollama on" wants one answer, not two fields to
reconcile.
codex_replay_reasoning is on by default and worth leaving on. It is writable so it can be
turned off from a console the day the route stops accepting a replayed reasoning blob, without
editing config.toml by hand on the serving machine.
The three write routes need --allow-admin, like the MCP login they mirror: they open a browser
on the serving host, delete a credential, and make an authenticated outbound call. The read half
is always mounted, so a console without admin can still show what is signed in and offer nothing
else.
When the config file will not load
GET /api/config describes the config in force, which is not always the file on disk. A save
that does not parse, or that sets a value Leviath refuses, leaves the last config that loaded in
force: runs keep starting on it, and nothing about them changes. Without a way to say so, that
looked exactly like an edit nobody had made.
config_error is how the server says so. It is absent while the file loads, and present with the
whole story while it does not:
{
"default_provider": "anthropic",
"config_mtime": 1756000000,
"config_error": {
"kind": "parse",
"path": "/Users/you/.leviath/config.toml",
"message": "expected `.`, `=`",
"line": 12,
"column": 8,
"since": 1756000420,
"note": "The running config is the last one that loaded; edits to this file take effect only once it parses again."
}
}kind is parse for a syntax error or a value of the wrong type. It is validation for a value
that parsed and was then refused, and read for a file that could not be read at all. A parse failure
carries line and column, both 1-based; a validation failure carries key instead, the dotted
config key it is about, such as model_providers.local. message is one line with no caret art in
it, ready to put in a banner. since is when this server first saw the file in this state, in unix
seconds.
config_mtime is always there, error or not: it is the mtime of the config in force, so a client
that just wrote the file can tell whether the write was picked up. While config_error is set it
names the last good save rather than the file on disk.
Nothing has to be restarted. Fix the file, and the next request answers with no config_error and
the new values. Announced as the config.health capability; a server that does not announce it
omits the field whether or not the file loads, so its absence proves nothing there.
PUT /api/config is checked before it writes, with the same rules the loader applies, so this API
cannot be the thing that breaks the file. A body that would produce a config this build refuses to
read back answers 400, and the file is left byte for byte as it was.
Override and fallback models
override_model on GET /api/config is the model every stage that allows a user default starts on
while it is set, ahead of what its blueprint names. fallback_model is the model a stage falls back
to when none of the models it names is configured on that machine, never ahead of them. Both are
bare model ids on default_provider. Both are always sent, null included, and that is the point:
a server that omits a key predates the field, which is a different answer from "nothing is set".
Read the absence as "cannot say" and show that, rather than showing an empty picker over a machine
that has a model pinned. Servers before 0.6 sent default_model instead, which behaved as
override_model does.
On PUT /api/config each key has three states, which the other fields in that body do not have:
| Body | What happens |
|---|---|
no override_model key |
the setting is left exactly as it was |
"override_model": null |
the setting is written away |
"override_model": "gpt-5" |
the setting is pinned to gpt-5 |
fallback_model reads the same way, and so do the seven provider keys (anthropic_key,
openai_key, google_key, xai_key, meta_key, openrouter_key, bedrock_key). null clears
a key, which takes the provider out of this install the way the setup wizard's remove does. An
empty string is a 400. Clearing override_model matters as much as setting it. A pinned model
runs every stage of every blueprint on that one model, and the cheap stages then pay a top-tier
price. Unset is the state most machines want: see
which entry a stage starts on. Sending null is
the only way to get back there through this API, in the same way remove_gateways is the only way
to delete a gateway.
"override_model": "" (or "fallback_model": "") is a 400, not a clear. An empty string is not
a model id, and a form that posts its empty box should be told rather than quietly lose the
setting. Nothing is written when it is refused.
default_provider takes no null. It is not optional in config.toml, because a machine always
has one, defaulting to anthropic. There is no unset state to write, and sending a different name
is the whole vocabulary.
Gateways
gateways on GET /api/config lists every [model_providers.<name>] entry, name-sorted, and
each one says what backs it: kind is script for a Rhai provider or openai-compatible for a
server that speaks OpenAI's chat API. Beside name, base_url, has_api_key and script, an
endpoint reports header_names and models. header_names holds the names of its extra headers,
never their values, because a header is where a second credential goes. models holds the ids it
falls back to when its server will not list them. extra_keys names a script's forwarded keys
the same way.
PUT /api/config takes the same fields on each gateway in gateways: kind, base_url,
api_key, script, headers (a name-to-value map) and models. Every field is optional and an
absent one leaves what the entry already had, so a console can edit a URL without knowing the
key or sending the headers back. An unknown kind, or an openai-compatible gateway with no
base_url, is refused with a 400 and nothing is written.
POST /api/models/probe is for the form before the write: it sends GET /models to base_url
with the given api_key and headers, exactly as the gateway would, and answers
{"models": [ids]} sorted. A server that refuses or does not answer is a 502 whose error is the
server's own text, and a base_url with no scheme is a 400. It is mounted only with
--allow-admin, like the write it precedes: it makes the serving host open a connection to any
address the caller names.
Each entry from GET /api/models also carries limits_source: api when the provider reported the
token limits itself, builtin when this build matched them off the model's name, and override
when a [model_capabilities] entry set them. Read it before treating a window as a fact. A
builtin figure for a model the table does not know is a guess, and region budgets resolve against
it. Beside it: supports_temperature and supports_tools. learned is true when the provider's
own listing described the model, and false for a row from this build's table. When the listing
carries them, released (Unix seconds), retires (the date the provider published) and pricing
(USD per million tokens: input_per_mtok, cached_input_per_mtok, cache_write_per_mtok,
output_per_mtok) come too, each null otherwise. Two lists say what the model takes and hands back:
input_types and output_types, mime type patterns such as text/*, image/* or
application/pdf, from this build's table corrected by the provider's listing and by
[model_capabilities]. A stage holding an image picks a model whose input_types cover it; see
More than text. Which providers can report what, and from where, is in
where a model's capabilities come from.
That is what lets a console show the catalog without fetching and re-parsing every script. No other
kind carries the key at all.
Validation checks more than syntax here. A provider needs initialize(config) and
inference(state, request), and a script defining only the first used to compile, initialize, cache
and then fail at the first inference, part-way into a run. POST /api/scripts/validate with
kind: "provider" answers that before the file is saved, and the loader refuses the same script
rather than accepting one the API called invalid. Nothing runs during validation: initialize is
read off the compiled AST, never called.
GET returns the source verbatim. A provider's key comes from initialize(config), which is the
[model_providers.<name>] table GET /api/config already reports as a boolean plus a list of key
names, or from env_var, which reads the daemon's environment. Neither value is in the file, so
there is nothing here for redaction to protect, and an editor that saved what it was shown would
write the redaction back over the real script.
Check scripts.providers in the capabilities list before offering the kind.
Feature detection
GET /api/config reports api_version, a capabilities list, and the server's limits. Check
those instead of calling a route and treating a 404 as "unsupported": a 404 also means "no such
run", and it costs a round trip per feature. The limits matter as much as the capability names: they
are where the page cap, file cap and listing cap actually live, so a client never has to hardcode
one.
runs.delete and runs.delete.bulk are the ones worth checking before drawing a button rather than
after clicking it. Finding out whether the other routes exist costs a wasted request; finding out
this way costs a deleted run.
What each capability means
Every string this server can announce, and the thing it promises. A server that omits one is older than that feature, not broken.
| Capability | The server has |
|---|---|
runs.envelope |
GET /api/runs answering {items, next_cursor, total, server_time} rather than a bare array |
runs.cursor |
Keyset paging on that route: cursor= in, next_cursor out |
runs.search |
q=, a case-insensitive substring over the run listing |
runs.search.context |
q_in=context, searching each run's current context window |
runs.search.logs |
q_in=logs, searching the tail of each stage's logs |
runs.search.journal |
q_in=journal, searching the whole run journal |
runs.fields |
fields=, trimming each item to the named top-level fields |
runs.ids |
ids=a,b,c, fetching exactly those runs in one request |
runs.since |
since=, filtering on whichever timestamp sort names |
runs.parent |
parent=none / parent=<run_id>. See listing by place in the tree |
runs.files.listing |
GET /api/agents/{id}/files, the run's own record of what it changed |
runs.files.workdir |
source=workdir on that route, reading the filesystem a directory at a time |
runs.files.mime_type |
mime_type on every file-listing entry, typed by the run's registry from the file's name |
models.mime_types |
input_types and output_types on every GET /api/models entry: the mime type patterns a model takes and hands back |
models.cached |
GET /api/models answers from a catalogue the server keeps. See the model catalogue |
spawn.parts |
parts and multipart/form-data on POST /api/agents. See attaching files |
messages.parts |
The same on POST /api/agents/{id}/message |
runs.blobs |
GET /api/agents/{id}/blobs and /blobs/{sha256}: the stored parts a run holds and their bytes. See a run's parts |
runs.files.raw |
GET /api/agents/{id}/files/raw?path=, a workdir file's bytes under its own content type |
runs.artifacts |
GET /api/agents/{id}/artifacts/{name}, the bytes of one file the run handed back, from the blob store or the workdir |
runs.result.artifacts |
artifacts on a run's answer as { name, path, mime_type, size, sha256 } objects rather than paths |
mime.registry |
GET /api/mime, the effective mime registry with each row's source |
mime.write |
PUT /api/mime and DELETE /api/mime, admin-gated, write a row into mime_types.toml or take one out |
runs.stages |
GET /api/agents/{id}/stages, the per-stage ledger |
runs.stages.cost |
cost_usd, unpriced_calls and cost_is_exact on each stage record, and the visits split beneath them |
runs.stages.models |
models on each stage record: the provider and model pairs that stage ran on, oldest first |
runs.waiting_on |
wait_reason on a run, saying what a parked run is parked on |
runs.delete |
DELETE /api/runs/{id}, which removes the record rather than cancelling the run |
runs.delete.bulk |
DELETE /api/runs with before or ids, bounded by max_ids |
logs.stage |
?stage= on the logs route: an index, or all |
logs.stream |
?stream= on it: output or logs |
context.history.page |
Paging on GET /api/agents/{id}/context/history |
context.region_kinds |
Region kinds spelled as the blueprint spells them. See region kinds |
events.waiting_on |
wait_reason on the socket too, not only on the run |
events.stage_and_tool |
stage_transition, tool_call_started and tool_call_finished as flat frames instead of the old world envelope |
events.spawn_parent |
parent_id on agent_spawned, placing a sub-agent in the tree the moment it starts |
events.title |
The run_renamed frame, plus title on every agent_status |
events.run_status |
One status vocabulary across the whole API. See statuses |
events.spend |
The agent_spend frame, sent as a run passes a figure in [limits] notify_spend_usd |
runs.cost |
cost_usd and subtree_cost_usd on the agent tree routes |
blueprints.envelope |
The paginated envelope on the blueprint listing |
blueprints.query |
q= on that listing |
blueprints.manifest |
The manifest itself on the blueprint detail route |
blueprints.validate.name |
POST /api/blueprints/validate accepting an installed name, not only a body |
blueprints.fan_outs |
fan_outs on the detail route. See fan-out limits |
blueprints.stage_routing |
stage_routing on the detail route: output_routing and context.reset per stage. See stage routing |
tools.list |
GET /api/tools?agent=, what an agent here can actually call |
update.plan |
GET /api/update, how this copy was installed and the command that upgrades it. See asking how to upgrade |
update.apply |
POST /api/update and GET /api/update/jobs/{id}, carrying that plan out. See pressing the button |
scripts.read |
The GET half of the scripts routes |
scripts.write |
That this build serves the write half |
scripts.providers |
provider as a fifth script kind, the machine's drop-in model providers |
scripts.candidates |
?include=candidates on the script listing, plus relative_path and declared on every entry |
scripts.mime_checks |
mime_check as a sixth script kind: the byte checks mime rows name |
config.gateways |
gateways on GET /api/config, the custom providers this machine has |
config.gateways.kinds |
kind, header_names and models on each gateway, and kind, headers and models accepted by PUT /api/config |
models.probe |
POST /api/models/probe, which asks an OpenAI-compatible server what it serves before a gateway for it is written; admin only |
fs.mkdir |
POST /api/fs/dirs, so a folder picker can offer "New Folder" rather than one that 404s |
interaction.feedback |
feedback beside approved: false on POST /api/agents/{id}/interaction. See answering a question |
providers.signin |
GET /api/providers and the three admin routes under it. See signing in to a subscription provider |
providers.quota |
?quota=true on GET /api/providers, and the quota object it adds. See subscription usage |
graphql |
POST /graphql, the GraphQL API beside these routes |
graphql.subscriptions |
GET /ws/graphql, the live frames with server-side filtering |
graphql.executions |
RunOutput.executions: what a run tried, with each call typed |
graphql.interactions |
RunOutput.interactions: every question a run put to a person, with the settlement typed |
graphql.inferences |
RunOutput.inferences: every trip a run made to a provider, retries and failovers included |
graphql.context_changes |
RunOutput.contextChanges: why each of a run's regions changed, beside the snapshots |
bytes.signed_urls |
Short-lived exp/sig links on the byte routes, minted by the GraphQL API |
runs.blueprint_snapshot |
blueprint_digest on every run, and the manifest copy each run keeps |
runs.export |
startRunExport and GET /api/exports/{id}, the whole store as one file. See below |
config.health |
config_error and config_mtime on GET /api/config, and the config_health frame on the socket. See below |
A few of those promises carry a consequence worth spelling out.
models.cached also covers X-Leviath-Catalog-Age and X-Leviath-Catalog-Complete on the
response, and ?refresh=1 to ask the providers again. It is safe to call when a page opens.
spawn.parts also covers @path tokens in task and region text, resolved inside the working
directory.
Without runs.stages.cost, a stage record carries tokens and no price, and the missing field is
not a zero. Without runs.stages.models, no stage record says what it ran on, which is not the
same as a stage that has yet to run. scripts.mime_checks puts the byte checks beside the config for the operator's rows,
and beside the agent for a blueprint's. config.gateways.kinds is how you tell that a gateway can
be an OpenAI-compatible endpoint rather than a script.
update.apply and scripts.write say what this build serves, not what this daemon mounts. That
is --allow-admin, which you find out by calling one of the routes and reading the status.
interaction.feedback also covers the "Deny with feedback" option on a tool approval. An older
daemon drops the field without a word, so offer the box only where this is announced.
providers.signin is the browser sign-in for a provider that has no API key. Without it a console
can write codex_enabled and has no way to complete the sign-in. That leaves the user enabled and
unable to run anything.
providers.quota also covers X-Leviath-Quota-Age and X-Leviath-Quota-Complete on the answer,
and ?refresh=1 to read the accounts again. Without it a missing quota says two things at once:
a daemon that ignored the parameter, or a subscription with nothing to report.
Without config.health, a missing config_error means nothing, so a console cannot tell a file
that loads from a daemon that would not say.
Writing a mime row
A custom mime type is where the registry earns its keep. A family, a
token rule, extensions, a magic prefix and a byte check turn a format Leviath has never heard of
into one it types, sizes and validates, rather than one that behaves like
application/octet-stream. GET /api/mime shows what is there; these two writes change it,
without leaving the browser for the config file.
PUT /api/mime adds a row or updates the one already there, writing mime_types.toml beside the
config, the same file lev mime add writes and the operator's own rows live in:
PUT /api/mime
{
"mime_type": "application/x-acme-scene",
"family": "model",
"text": false,
"tokens": { "per_pixel": 750, "max": 1600 },
"extensions": ["scene"],
"magic": "41434D45",
"stand_in": "[{type} {size}] {name}",
"check": "checks/scene.rhai"
}Every field but mime_type is optional, and only the fields sent are changed, so a later PUT
that carries just {"mime_type": "...", "extensions": [...]} adds an extension and leaves the
rest. tokens is one of { per_byte }, { per_pixel, max? }, { per_second }, { per_page } or
{ fixed }.
The answer is {"mime_type", "created"}, where created is false when the row was already there.
The row is validated the way lev mime add validates it, before anything is written. A 400 comes
back for a type that is not type/subtype, a token rule naming none or more than one rate, a
magic that is not hex, or a check script that will not compile. The file is untouched.
DELETE /api/mime?mime_type=<type> takes a row out; a type with no row of its own there is a 404.
Both need --allow-admin, and both are announced as mime.write, so a console offers "New
type…" where it will land and hands over the TOML to paste where it will not.
Live updates over WebSocket
Connect to /ws (all agents) or /ws/agents/{id} (one run) with ?token=<t>; the server streams
ServerEvent frames as the run progresses:
sequenceDiagram
participant Browser
participant Serve as lev serve
Browser->>Serve: GET /ws/agents/{id}?token=…
Serve-->>Browser: 101 Switching Protocols
loop while the run is live
Serve-->>Browser: {stage changed}
Serve-->>Browser: {tokens updated}
Serve-->>Browser: {awaiting input}
end
Serve-->>Browser: {done}
The frames
Every frame is a JSON object with a type. Every frame except daemon_link and config_health
carries a run_id, which is what /ws/agents/{id} filters on. daemon_link reaches every
subscription including a per-run one, because it explains why a run's frames stopped;
config_health reaches only /ws, because a run in flight keeps going on the config it started
with either way.
type |
Sent when | Beyond agent_id and run_id |
|---|---|---|
agent_spawned |
A run first appears | blueprint, and parent_id for a sub-agent |
agent_status |
Status, stage, iteration or tool count moves | status, stage, iteration, tool_calls, accepts_messages, wait_reason, title |
run_renamed |
The run acquires a generated title | title |
tokens |
The run's token totals move | prompt_tokens, completion_tokens, cached_tokens, cache_write_tokens |
context_update |
The context window's usage moves | total_tokens, max_tokens |
stage_transition |
A new stage is entered | from, to, iteration |
tool_call_started |
A tool call goes to the async lane | call_id, tool |
tool_call_finished |
That call returns | call_id, tool, ok, summary |
log |
A log or output line is written | line |
agent_spend |
The run's spend passes a figure named in [limits] notify_spend_usd |
threshold_usd, total_usd, complete, stage |
interaction_needed |
The run is blocked on a person | request |
agent_completed |
The run reaches a terminal status | status, result (its error), final_output |
daemon_link |
This server's link to the daemon changes | connected, daemon, restarted, restart_advised |
config_health |
config.toml stops loading, loads again, or breaks for a different reason |
healthy, path, error, config_mtime |
agent_spend arrives while the run is still going, which is the point: a run that quietly spends
far more than intended looks, from the outside, exactly like one making ordinary progress. Each
figure in [limits] notify_spend_usd is announced once per run, the first time the total passes it,
and stage names the stage that was running when it crossed. Nothing is emitted for an operator who
has not listed any figures.
complete says whether every call behind total_usd could be priced. When it is false the run has
spent at least that much and more by an unknown amount. It is a different question from whether the
priced part came from the provider's own figures or was reconstructed from published rate cards.
That is what cost_is_exact on the run record answers, so a total can be complete and still be a
reconstruction.
A run is created untitled and named a moment later, once a model has shortened its prompt into a
title. run_renamed is that moment. The same title then rides every agent_status frame, so a
client that connected or reconnected after the rename reads the name off the next status instead of
fetching the run. Both are the events.title capability; without it a client has to poll each new
run until it has a name. title is absent, not null, while a run has none.
Naming can also fail. The call retries a transient refusal and then walks the run's own model
candidates, the same chain its stage inference fails over along, so one provider being unreachable
no longer costs the run its name. When every candidate is spent, title_error on the run says what
stopped it, and stays null while titling is still under way. Poll that field rather than waiting
forever on a run_renamed frame that is not coming.
status, on agent_status and on agent_completed, is the word GET /api/runs uses for the same
run. See Statuses for the list and for what a server that predates
events.run_status sends instead.
wait_reason is present only on a parked run, and says what it is parked on rather than making
you fetch the run to find out. A run Leviath paused carries
{"reason":"needs_setup","blocker":"...","remedy":"..."}, where remedy is a sentence with the
provider's own error in it and blocker is one of provider_missing, credits_exhausted,
auth_failed, forbidden, provider_unreachable, provider_timed_out, provider_failed or
providers_unavailable. Troubleshooting
says what each one asks of you. ok on tool_call_finished is false for a result the engine
refused or could not run, so a client should not read a finish frame as a success on its own.
stage_transition, tool_call_started and tool_call_finished used to arrive wrapped as
{"type":"world","event":{…}}. They are flat frames of their own as of API version 0.4.0,
announced as the events.stage_and_tool capability on GET /api/config; parent_id on
agent_spawned is events.spawn_parent. There is no longer a world frame.
When the daemon restarts
The stream stays open across a daemon restart. lev serve reconnects to the daemon
on its own, so your socket never has to. What you see is one daemon_link frame when the daemon's
events stop, and one when they resume:
{"type":"daemon_link","connected":false,"daemon":{"version":"0.4.0","build":"3ba95219","pid":4242},"restarted":false}
{"type":"daemon_link","connected":true,"daemon":{"version":"0.4.0","build":"3ba95219","pid":4301},"restarted":true}restarted says whether the daemon that came back is a different process from the one before.
daemon is absent until the daemon has introduced itself, which every current daemon does on
connect.
The daemon can come back on a different build than the running lev serve. The usual cause is a
lev update with the server left running. The frame then also carries restart_advised, a
sentence that names both builds and says to restart lev serve. Every subscriber that connects
while that is true, or while the daemon is unreachable, gets a daemon_link frame first thing. A healthy
stream sends none, so a client that ignores the type sees exactly what it always saw.
Requests keep working across a version gap as long as the two ends still understand each other. A
request that fails because they no longer do answers 502 with the same sentence, where a
daemon that is not answering at all is a 503. Retrying helps the second; only restarting
lev serve helps the first.
Asking for a shape
Add output_format to ask for the answer in a particular shape. Any label works, because nothing
converts between shapes: the label reaches the model, which produces the bytes.
curl -X POST http://localhost:3000/api/agents \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"blueprint":"reviewer","task":"Review the auth module",
"output_format":"a2ui",
"output_instructions":"One card per finding, highest severity first."}'Then read it back from GET /api/agents/{id}/result, where final_output carries the answer, its
format label, and the stage that produced it. Add output_schema when you want the answer validated
against a JSON Schema. A format that differs from the blueprint's retires any Rhai validator and
JSON schema the blueprint declared, and the spawn response says so: a warnings array beside the
run id names each retired check. Final outputs covers the whole cascade.
Spawning with a signed webhook
curl -X POST http://localhost:3000/api/agents \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"blueprint":"coder","task":"Add input validation",
"callback_url":"https://example.com/hook","callback_secret":"whsec_…"}'Four things to know about the delivery.
It carries the answer. final_output holds whatever the agent submitted, so your receiver
learns what the run concluded without a second request. The result field beside it is the run's
error, which is what it has always been. See Final outputs.
It is signed. Verify the X-Leviath-Signature: sha256=<hex> header against your
callback_secret before trusting the body. A callback_secret without a
callback_url is refused with a 400: the secret signs the callback body, so a
request that sends one and no URL is asking for a signed callback that can never
fire.
It carries a stable delivery_id, of the form agent_completed:<run_id>, in both the signed
body and the X-Leviath-Delivery header. Stable is the important word: a retried attempt, and a
completion re-fired after a daemon restart, both send the same id. So your receiver can deduplicate
with a plain key check and handle each completion exactly once.
It retries on transient failures, meaning network errors, timeouts, 5xx, 429, and 408, with exponential backoff. Every field below has a safe default, so you can leave the block out entirely:
[webhook]
max_retries = 3 # retries after the first attempt; 0 disables retries
base_delay_ms = 500 # first backoff; doubles per retry
max_delay_ms = 30000 # cap on any single backoff
timeout_secs = 10 # per-attempt request timeoutTip
The Lair is a full reference client for this API (connection, spawn, live dashboard, blueprint editing, MCP and policy management), built on the same typed endpoints.