voyagervygr
Source
AbstractRust · single static binaryMIT

Deep research that leaves a paper trail.

Voyager is a configurable deep-research CLI. It searches the web through provider chains (keyless DuckDuckGo out of the box, seven more behind env keys), fetches and reduces pages to text, and runs a plan-search-synthesize loop over any LLM backend: the model already configured in your agent harness, a local Ollama, or any OpenAI-compatible provider.

Install vygrRead the source$cargo install vygr

8 providers · 4 LLM backends · MCP server · evidence on disk

Fig. 1run artifacts · agents/voyager/<run>/
  • prompt.mdthe question + run parameters
  • plan.jsonsub-queries and chosen depth
  • reflections.jsondistilled notes per level
  • sources.jsonscored sources, with content
  • answer.mdreport with [n] citations
Every research run writes its full evidence trail to disk. Inspect it, diff it, cite it.
§ 1

Method

The loop

vygr research runs a real iterative loop: plan the sub-queries, then repeat search, fetch, score and reflect until the model has enough evidence to synthesize a cited report. Nothing is hidden; every stage leaves an artifact.

  1. 01--depth 2..4

    Plan

    An LLM turns the question into sub-queries and picks a depth within the range you allow.

  2. 02--all

    Search

    Each level goes back to the web through the provider chain, with concurrent fan-out and URL dedup on request.

  3. 03vygr extract

    Fetch

    Top results are fetched over HTTP and reduced to structured, readable text, boilerplate dropped.

  4. 04BM25

    Score

    BM25 ranks the collected evidence so only what matters enters the context window.

  5. 05reflections.json

    Reflect

    The LLM distills notes and follow-up queries from the gaps it finds. Breadth halves every level.

  6. 06answer.md

    Synthesize

    A markdown report with numbered citations, written only after the budget guard approves.

§ 2

Materials

Search providers

Search is pluggable at the SearchProvider seam. DuckDuckGo works with zero configuration; the rest wake up when their environment key appears. Any of them compose into chains, and every backend is rate-limited, retried and cached behind the scenes.

Table 1provider registry
ProviderAccess
ddgskeyless
braveBRAVE_API_KEY
tavilyTAVILY_API_KEY
exaEXA_API_KEY
serperSERPER_API_KEY
jinaJINA_API_KEY
kagiKAGI_API_KEY
searxngkeyless

Table 1 · provider registry, 8 backends

Fallback chain

default_provider = "ddgs,brave"

The first provider with a non-empty result set wins; failures degrade gracefully to the next link.

Concurrent fan-out

vygr search "q" --all

Searches every provider at once, merges and deduplicates by normalized URL, and records which provider returned each hit.

Cache and politeness

query-class TTLs

News, standard and reference queries get TTLs of 5 min, 1 h and 24 h. Per-provider rate limiting with retry and backoff, --no-cache to opt out.

§ 3

Results

Evidence you can audit

A report from a black box is a rumor. Voyager treats a research run like a lab notebook: the plan, the sources, the reflections and the synthesis are all written down, in plain files you can read.

Listing 1a research run
$ vygr research "state of WebAssembly GC in 2026"
--llm pi --depth 2..4--budget-usd 0.50
plan 3 sub-queries · depth 3
level 1 12 hits · 5 fetched · BM25 kept 4
level 2 breadth 2 · 2 follow-up queries
synthesize answer.md · 9 citations
cost $0.213 of $0.50
→ artifacts in agents/voyager/20261002-0914-wasm-gc/
Illustrative trace. Diagnostics go to stderr; with --format json the same data arrives on stdout.
  1. i

    Artifacts on disk

    prompt.md, plan.json, reflections.json, sources.json and answer.md under agents/voyager/<run>/. The report cites the sources it actually used.

  2. ii

    Structured answers

    Pass --output-schema and synthesis switches to JSON-constrained output with parse-level validation, ready for the next program in your pipeline.

  3. iii

    Honest accounting

    Every LLM call is priced from the models.dev catalog and checked against --budget-usd before it runs. Reports carry the final cost_usd.

§ 4

Apparatus

The command surface

Eleven subcommands, one contract. Everything a human reads in a terminal, an agent can read as JSON, so the same tool serves both.

Table 2command surface
vygr search <q>Provider-chain search: --all fan-out, --extract-top N, time ranges, domain filters, cache control
vygr extract <urls>Fetch pages and reduce them to plain, structured text
vygr research <q>The full loop: plan, levels, reflections, synthesis; artifacts under agents/voyager/<run>/
vygr plan <q>Offline preflight of a research run
vygr providersSearch providers, their env keys and readiness
vygr models [provider]Browse the models.dev catalog and pricing
vygr serveMCP server over stdio: search, extract, research, get_artifact
vygr schemaMachine-readable self-description, written for agents
vygr configEffective configuration and file paths
vygr init --agent piInstall the agent skill into your harness
vygr cache dir|clearInspect or clear the search cache
§ 5

Backends

Bring your own model

The loop is model-agnostic: everything pluggable hangs off a singleLlmClient seam, and the same run reads the same whether the brains are in the cloud, on your laptop, or inside your coding agent.

--llm pi--llm claude--llm codex

Reuse the harness

vygr shells out to the agent harness in print mode and pipes the prompt through stdin. The research loop runs on the model and credentials you already have configured. No MCP pass-through, no duplicate API keys.

--llm ollama:qwen3:8b

Local and cheap

A native Ollama client with thinking control handled for you. Reasoning models that answer through <think> blocks are stripped back to the answer, and --budget-usd caps what a run may spend.

--llm openrouter:deepseek/deepseek-chat

Any provider, priced

Any OpenAI-compatible provider from the models.dev catalog. vygr models browses models and pricing; the catalog (cached 24 h) also feeds the cost accounting of every run.

MCP server included

vygr serve exposes search, extract, research and get_artifact over stdio, with token-safe paged reads of run artifacts. Register it with your harness, install the bundled skill, or pull it straight from the skills registry:

$ pi mcp add voyager -- vygr serve
$ vygr init --agent pi|claude-code|codex|cursor|generic
$ npx skills add jaltez/voyager
§ 6

Distribution

Install

One Rust workspace, one static binary. Nothing below the CLI knows that clap exists, and nothing above it needs to know Rust.

crates.io

cargo install vygr

The released binary. Installs the vygr command; workspace crates come along.

GitHub releases

./install.sh

Prebuilt tarballs for linux and macOS, x64 and arm64. Falls back to cargo when no binary matches.

From a checkout

cargo install --path crates/cli --bin vygr

Build from source against your own toolchain.

workspace crates: vygr-core · vygr-providers · vygr-llm · vygr-research