Reuse the harness
vygr shells out to the agent harness in print mode and pipes the prompt through stdin. The research loop runs on the model and credentials you already have configured. No MCP pass-through, no duplicate API keys.
Voyager is a configurable deep-research CLI. It searches the web through provider chains (keyless DuckDuckGo out of the box, seven more behind env keys), fetches and reduces pages to text, and runs a plan-search-synthesize loop over any LLM backend: the model already configured in your agent harness, a local Ollama, or any OpenAI-compatible provider.
8 providers · 4 LLM backends · MCP server · evidence on disk
Method
vygr research runs a real iterative loop: plan the sub-queries, then repeat search, fetch, score and reflect until the model has enough evidence to synthesize a cited report. Nothing is hidden; every stage leaves an artifact.
An LLM turns the question into sub-queries and picks a depth within the range you allow.
Each level goes back to the web through the provider chain, with concurrent fan-out and URL dedup on request.
Top results are fetched over HTTP and reduced to structured, readable text, boilerplate dropped.
BM25 ranks the collected evidence so only what matters enters the context window.
The LLM distills notes and follow-up queries from the gaps it finds. Breadth halves every level.
A markdown report with numbered citations, written only after the budget guard approves.
Materials
Search is pluggable at the SearchProvider seam. DuckDuckGo works with zero configuration; the rest wake up when their environment key appears. Any of them compose into chains, and every backend is rate-limited, retried and cached behind the scenes.
| Provider | Access | Notes |
|---|---|---|
| ddgs | keyless | Keyless out of the box, the default first link |
| brave | BRAVE_API_KEY | Native freshness filter for time ranges |
| tavily | TAVILY_API_KEY | Native time-range and domain filters |
| exa | EXA_API_KEY | |
| serper | SERPER_API_KEY | |
| jina | JINA_API_KEY | Markdown responses, tolerant parser |
| kagi | KAGI_API_KEY | |
| searxng | keyless | Self-hosted instance, keyless |
Table 1 · provider registry, 8 backends
default_provider = "ddgs,brave"
The first provider with a non-empty result set wins; failures degrade gracefully to the next link.
vygr search "q" --all
Searches every provider at once, merges and deduplicates by normalized URL, and records which provider returned each hit.
query-class TTLs
News, standard and reference queries get TTLs of 5 min, 1 h and 24 h. Per-provider rate limiting with retry and backoff, --no-cache to opt out.
Results
A report from a black box is a rumor. Voyager treats a research run like a lab notebook: the plan, the sources, the reflections and the synthesis are all written down, in plain files you can read.
prompt.md, plan.json, reflections.json, sources.json and answer.md under agents/voyager/<run>/. The report cites the sources it actually used.
Pass --output-schema and synthesis switches to JSON-constrained output with parse-level validation, ready for the next program in your pipeline.
Every LLM call is priced from the models.dev catalog and checked against --budget-usd before it runs. Reports carry the final cost_usd.
Apparatus
Eleven subcommands, one contract. Everything a human reads in a terminal, an agent can read as JSON, so the same tool serves both.
| vygr search <q> | Provider-chain search: --all fan-out, --extract-top N, time ranges, domain filters, cache control |
| vygr extract <urls> | Fetch pages and reduce them to plain, structured text |
| vygr research <q> | The full loop: plan, levels, reflections, synthesis; artifacts under agents/voyager/<run>/ |
| vygr plan <q> | Offline preflight of a research run |
| vygr providers | Search providers, their env keys and readiness |
| vygr models [provider] | Browse the models.dev catalog and pricing |
| vygr serve | MCP server over stdio: search, extract, research, get_artifact |
| vygr schema | Machine-readable self-description, written for agents |
| vygr config | Effective configuration and file paths |
| vygr init --agent pi | Install the agent skill into your harness |
| vygr cache dir|clear | Inspect or clear the search cache |
Backends
The loop is model-agnostic: everything pluggable hangs off a singleLlmClient seam, and the same run reads the same whether the brains are in the cloud, on your laptop, or inside your coding agent.
vygr shells out to the agent harness in print mode and pipes the prompt through stdin. The research loop runs on the model and credentials you already have configured. No MCP pass-through, no duplicate API keys.
A native Ollama client with thinking control handled for you. Reasoning models that answer through <think> blocks are stripped back to the answer, and --budget-usd caps what a run may spend.
Any OpenAI-compatible provider from the models.dev catalog. vygr models browses models and pricing; the catalog (cached 24 h) also feeds the cost accounting of every run.
vygr serve exposes search, extract, research and get_artifact over stdio, with token-safe paged reads of run artifacts. Register it with your harness, install the bundled skill, or pull it straight from the skills registry:
Distribution
One Rust workspace, one static binary. Nothing below the CLI knows that clap exists, and nothing above it needs to know Rust.
cargo install vygrThe released binary. Installs the vygr command; workspace crates come along.
./install.shPrebuilt tarballs for linux and macOS, x64 and arm64. Falls back to cargo when no binary matches.
cargo install --path crates/cli --bin vygrBuild from source against your own toolchain.
workspace crates: vygr-core · vygr-providers · vygr-llm · vygr-research