Agent catalog
Leviath ships with seven pre-built agents. lev setup installs them into ~/.leviath/agents/
(scripting it? pass --install-agents), one directory per agent, each holding an agent.leviath
blueprint. Run any of them by name:
lev run coder --task "Build a CLI that converts CSV to JSON"Each entry names the models its stages prefer. Those are first choices, not requirements: every bundled agent lists all five providers, and a stage falls back to the one you configured.
This page is also a set of worked examples. Each section shows the agent's real stages and how
they route, so you can copy the patterns into your own blueprint (lev create my-agent scaffolds
one, then read Agents). The diagrams are simplified: they show each agent's main
path, and most draw the edge into its error-recovery stage, but a real graph has more edges than
these. lev validate <agent> prints every one of them.
Tip
Pick by the shape of the work: a codebase change (the coding agents), a question to answer from
sources (the research agents), or a recurring chore like triaging logs. Not sure? Run coder.
coder aside, every agent that has more than one thing to cover
fans out: data-analyst, deep-researcher, log-analyzer, reviewer,
and wide-researcher all work on several at once instead of one after another.
coder
Discover the repo, plan with your sign-off, optionally spike an uncertain approach, implement, and review. The coding agent: reach for it for any change to a codebase.
flowchart TD
discover --> plan
plan -->|revise| plan
plan --> prototype
plan --> implement
prototype --> implement
prototype -->|re-plan| plan
prototype -->|stuck| reassess
implement --> review
implement -->|re-plan| plan
review -->|issues| implement
implement -->|stuck| reassess
reassess --> implement
implement -->|error| error_recovery
error_recovery --> implement
lev run coder --task "Add rate limiting to the public API"discover answers what the repository is and how to verify work in it before anything is planned,
so the plan is grounded in the project rather than in guesswork.
The plan stage runs in interactive_points mode, so it stops for your approval before any code
is written. That is the human-in-the-loop pattern. Unattended, the checkpoint
resolves as approved rather than stranding the run, so --yolo in CI still works. Set
unattended = "ask" on the interaction point if you would rather an unattended run wait.
plan also chooses between going straight to implement and spiking first with prototype when
the approach is uncertain, and reassess is reached only on a stuck edge. discover and plan
run on Sonnet; implement, review, and reassess step up to Opus, so it is a
multi-model blueprint.
reviewer
Review only: a fast scan pass, then a deeper look at correctness, security, and architecture, ending in a ranked report. Reach for it to vet a diff or PR.
flowchart LR
discover --> scan
scan --> split_review
split_review -->|"fan out: one worker per area"| deep_review
deep_review --> report
deep_review -->|error| error_handler
error_handler --> deep_review
lev run reviewer --task "Review the changes on the feature/auth branch"
lev run reviewer --diff @change.patch --criteria "does the code produce what after.png shows?" \
--attach before.png:screenshots --attach after.png:screenshotsScreenshots are a typed input: the screenshots region accepts image/* and holds six, so a
mockup or a before-and-after reaches the model. A model that can see gets it as pixels, and any
other gets a line naming the file. A @path in --criteria attaches the same way. See
Mime.
The two-pass split is deliberate: scan runs on Sonnet to flag areas, then the review itself
escalates to Opus to scrutinize only what was flagged, which keeps the expensive model focused.
split_review is a fan-out stage: one worker per file, module, or group of
hunks, all reviewing at once. deep_review merges their findings, re-checks the blocking ones,
covers any area whose worker failed, and then looks for what no single area could show. A small
change is one work item, so a two-file diff does not pay for a fan-out it does not need.
data-analyst
Searches the web for data on a subject, builds a clean CSV of it, and hands back a summary of what the numbers say. Reach for it when you want a dataset you can open, not a paragraph about one.
flowchart TD
scope --> split
scope -.->|only if splitting is exhausted| build
split -->|fan out| gather_worker
gather_worker --> build
build --> present
present --> done["Summary + dataset.csv"]
lev run data-analyst --task "EV registrations by country, 2015 to 2024"
lev result <run-id> # what the numbers sayscope decides the table's columns before anything is gathered, which is what keeps a hundred rows
from drifting into a hundred shapes. split runs in fan_out mode, one worker per slice of the
subject, so a broad question is gathered in parallel rather than one source at a time. See
Sub-agents and fan-out.
Each worker hands back CSV rows through submit_output, and require_output makes
that a guarantee rather than a hope. The worker's format = "csv" label travels with the
submission, so the merge knows what shape it is being handed.
present is instructed to name data/dataset.csv in its artifacts, so a caller fetches the file
rather than parsing its path out of prose.
researcher
General-purpose research: gather, analyze, summarize, with a refinement loop. Reach for it for a quick, focused answer.
flowchart LR
gather --> analyze
analyze -->|need more| gather
analyze --> summarize
analyze -->|error| error_recovery
error_recovery --> analyze
lev run researcher --task "What changed in the HTTP/3 spec this year?"The analyze stage loops back to gather when a specific sub-topic is thin, then moves to
summarize once the picture holds. analyze runs on Opus; gather and summarize stay cheap, the
multi-model split again.
wide-researcher
Broad landscape survey: cast a wide net, compare approaches, deep-read the interesting threads, then write an overview with recommendations. Reach for it to map a whole space.
flowchart TD
survey --> investigate
investigate -->|"fan out: one researcher per thread"| compare
compare -->|gaps| survey
compare --> deep_dive
deep_dive --> compare
compare --> challenge
challenge -->|"needs more evidence"| survey
challenge --> summarize
summarize --> polish
polish --> summary
compare -->|error| error_recovery
error_recovery --> compare
lev run wide-researcher --task "Survey approaches to vector database indexing"investigate is a fan-out stage, and its workers are full researcher runs
rather than stages of this blueprint. Every thread the survey found is researched at the same time,
each with its own clean context window, and their findings merge into compare. A worker that finds
its thread is really several independent subjects can split again, one level further.
challenge and polish work as they do in deep-researcher. Every route to the
writing stage passes through an adversary that can send the survey back for more evidence. The
finished overview is rewritten in plain language, without any fact, number, citation or caveat
changing.
compare is then the hub: widen coverage (back to survey), pull one thread for a focused
deep_dive, or finish.
deep-researcher
Thorough single-topic investigation: follows citation chains, cross-checks claims, and produces a structured, cited report. Reach for it when rigor and sources matter.
flowchart TD
gather --> investigate
investigate -->|"fan out: one researcher per sub-question"| analyze
analyze -->|gaps| gather
analyze --> follow_citations
follow_citations --> analyze
analyze --> challenge
challenge -->|"needs more evidence"| gather
challenge --> synthesize
synthesize --> polish
polish --> summary
analyze -->|error| error_recovery
error_recovery --> analyze
lev run deep-researcher --task "Investigate the evidence for X causing Y"A thorough investigation is usually several questions wearing one coat. investigate is a
fan-out stage that splits them out and runs each as its own researcher sub-agent
in parallel, merging what comes back into analyze. A sub-researcher that finds its own slice is
several independent subjects can split again, one level further.
follow_citations is a dedicated targeted-read stage: analyze flags a specific cited source, the
stage pulls and reads it, then hands control back.
challenge is the gate to the report, and every path to the writing stage runs through it. A
different vendor's model attacks the analysis, marks what holds as well as what is weak, and can send
the run back to gather more evidence. A different vendor on purpose: an adversary that shares the
writer's blind spots is not an adversary. Making it optional does not work, because a model that
feels confident will not elect to be attacked, and confidence is the failure it exists to catch.
polish rewrites the finished report in plain language without touching a fact, number, citation or
caveat. It exists because the stage that gathers the most evidence is not the one that writes the
clearest prose, and asking one model for both gets a worse version of each.
Per stage, the models are chosen on measurement rather than by defaulting to one family. A fast broad-search model gathers, a cheaper reasoning model analyses, a strong writer synthesises, a different vendor challenges, and a plain-language model polishes. See providers for how a stage picks its model and falls back.
log-analyzer
Analyzes log files for anomalies, trends, and error patterns through a scripted analyze and script loop, keeping a severity-ranked findings index. Reach for it to triage a noisy log.
flowchart LR
ingest --> split_logs
split_logs -->|"fan out: one worker per file or window"| analyze
analyze --> script
script -->|refine| script
script --> analyze
analyze --> report
analyze -->|error| error_recovery
error_recovery --> analyze
lev run log-analyzer --task "Find the error patterns in /var/log/app.log"split_logs is a fan-out stage: one sweeper per log file, per service, or per
time window of a single large file, all reading at once. analyze merges what they found, and a
slice whose worker failed comes back as unswept rather than silently missing. One log file is one
work item, so a single file does not pay for a fan-out.
analyze (on Opus) then hands off to script to write and run parsing or aggregation code, which
can refine itself before returning results. Findings persist in a context region
across passes so the report ranks them by severity.
3D models (Meshy)
Four agents build 3D assets through the Meshy provider.
Configure Meshy (set MESHY_API_KEY, or lev setup and choose Meshy) and hand each one its input.
sprite-to-3dturns a sprite sheet or character image into a rigged, game-ready model: it draws clean multi-view references, critiques them against the source, builds the mesh, rigs it, and checks the result.lev run sprite-to-3d --task "the character on the left" --attach sheet.png:source.image-to-modelturns one image straight into a textured model, no reference-drawing step.lev run image-to-model --task "matte plastic" --attach toy.png:image. Pure Meshy: it needs no text provider at all.text-to-image-to-modeldraws a reference image from a description, then builds a model from it.lev run text-to-image-to-model --task "a stout clay teapot". The draw step needs an image-capable provider (OpenRouter's image model by default); the build needs only Meshy.model-to-animated-modelrigs an existing mesh and applies an animation from Meshy's library.lev run model-to-animated-model --task "run" --attach hero.glb:source_model. Pure Meshy, no text provider.
Each hands the finished model back as a model/gltf-binary artifact; lev result <run> --out .
writes it out. A single-provider agent like image-to-model or model-to-animated-model records
that mesh as its answer directly, with no submit_output turn, so a pure "bytes in, bytes out"
pipeline runs with only Meshy configured.
Running one
Every agent runs the same way. Name it and hand it a task:
lev run deep-researcher --task "Survey the state of solid-state batteries"To build your own, read how blueprints are structured in Agents. Multi-stage workflows covers how the stage graph routes and recovers, and Sub-agents and fan-out covers how the parallel agents split work.