Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

rewriter-queue

A resumable job queue for multi-agent LLM code synthesis. You submit a source tree and a mission; a small organization of agent roles (implementer, security/maintenance/production-readiness reviewers, a merge agent for large inputs) iterates on it — patching rather than rewriting where it can, checkpointing every real step so an interrupted run resumes instead of starting over — until the reviewers are satisfied or the iteration budget runs out.

Submit and watch jobs from a terminal, from Zed, or from Claude Code.

Not a hosted service — three small Rust binaries you run yourself, against whatever OpenAI-compatible inference endpoint you point them at.

Start with the tutorial.

Tutorial

This walks through the whole loop once, end to end: get a binary, point it at a model, submit something to rewrite, and watch it work. Fifteen minutes if you’ve already got an OpenAI-compatible inference server running somewhere; longer if you need to stand one up first (see Remote deployment for that part).

1. Get the binaries

Download the archive for your platform from the latest release and unpack it — see Installation if you’d rather build from source. Either way you end up with two binaries sitting next to each other: rewriter-queue and ai-org-orchestrator. They need to stay in the same directory — the queue’s worker finds the orchestrator by looking beside its own binary, not on PATH.

Put that directory on your PATH, or just remember the full path for the commands below.

2. Tell it where your model lives

Create ~/.config/rewriter/config.toml:

[[providers]]
name = "local"
type = "openai-compat"
base_url = "http://localhost:8001/v1"
models = ["your-model-id"]

Point base_url at any OpenAI-compatible /v1/chat/completions endpoint — a llama.cpp/vLLM server you’re running, or a hosted API. If you don’t have one yet, that’s the whole subject of Remote deployment.

3. Pick a queue mode

For trying this out on one machine, skip the server entirely — every command below falls back to a local, file-backed queue when there’s no [queue] section in the config. If you’re running the inference backend on a separate box and want to submit jobs from your laptop, start the server there instead and add a [queue] url = "..." pointing at it (see the root README’s quickstart for the server command and token setup).

This tutorial uses local mode — nothing extra to start.

4. Submit something

Pick a small, real piece of code you want rewritten or reviewed — a single source file or a small directory is a better first run than a whole project. The workspace directory holds the tool’s working state as it goes, plus one thing you provide before submitting: the mission.

The mission is a required artifact — a short statement of what V2 should actually do, not just “port this file.” Without it, the first agent in the pipeline has to guess your intent from raw source alone, which shows up later as a contract that’s technically correct but not what you meant. It’s stored the same way as every other artifact this pipeline produces (objective_contract.md, schema.rs, …), at <workspace>/artifacts/mission.md, and the run refuses to start without one:

mkdir -p /tmp/my-first-run/artifacts
cat > /tmp/my-first-run/artifacts/mission.md <<'EOF'
# Mission

V1 is a small in-memory counter store. Rebuild it as V2 with identical
behavior: increment a key's count, and read a key's current count (0 for
an unseen key).
EOF

rewriter-queue submit \
  --source /path/to/the/code \
  --workspace /tmp/my-first-run \
  --max-iter 10

That prints something like queued job #1 (0 ahead). In local mode with nothing else queued, it starts immediately.

5. Watch it work

rewriter-queue watch

A live view that refreshes once a second: which job is running, what stage it’s on (ObjectiveContract → test matrix → inductive analysis → schema → V2 synthesis), and how long it’s been going. Ctrl-C to stop watching — the job keeps running either way.

Want more detail on one job specifically?

rewriter-queue status 1

Shows the last 20 lines of its log, including which agent is running and the token counts for each call.

6. See what it’s produced, before it’s even done

You don’t have to wait for the job to finish to look at what it’s built so far:

rewriter-queue artifacts 1          # list what's been emitted
rewriter-queue artifact 1 objective_contract.md   # read one of them

objective_contract.md is usually the first thing worth reading — it’s the agreed-on shape of what’s being built (public API, invariants, edge cases, explicit exclusions) before any code gets written.

7. Get the result

In local mode there’s nothing to download — the workspace directory you passed to submit is where everything lands, no copy needed:

ls /tmp/my-first-run/artifacts/v2/   # the actual synthesized code

(rewriter-queue download 1 --out ./result is for the remote-server case — see Remote deployment — where the job ran on a different machine and you need its artifacts/ tree copied to yours in one call instead of one file at a time.)

8. If it gets interrupted

Kill rewriter-queue serve (or just Ctrl-C a local-mode run) mid-job and resubmit isn’t what you want — restart the server (or just try watch again in local mode) and the same job picks back up from wherever it had already checkpointed, not from the beginning. This is true down to the level of individual agent calls, not just whole pipeline stages — see the root README’s design notes if you’re curious how.

What’s next

  • Installation — the binary download details, and building from source if you’d rather.
  • Zed setup — submit and watch jobs from Zed’s agent panel instead of a terminal.
  • The Claude Code plugin — the same queue tools inside Claude Code, plus each reviewer persona as a standalone slash command.
  • Remote deployment — running the inference backend and the queue server on a separate (e.g. GPU) machine from the one you’re submitting jobs from.

Installation

Prebuilt binaries

Every release ships prebuilt archives for:

PlatformArchive
Linux x86_64rewriter-queue-x86_64-unknown-linux-gnu.tar.gz
Linux aarch64 (ARM64)rewriter-queue-aarch64-unknown-linux-gnu.tar.gz
macOS Intelrewriter-queue-x86_64-apple-darwin.tar.gz
macOS Apple Siliconrewriter-queue-aarch64-apple-darwin.tar.gz
curl -LO https://github.com/AndyGauge/rewriter-queue/releases/latest/download/rewriter-queue-<your-target>.tar.gz
tar xzf rewriter-queue-<your-target>.tar.gz

Each archive contains two binaries — rewriter-queue and ai-org-orchestrator — plus the license files. Keep both binaries in the same directory. The queue’s worker finds the orchestrator by looking beside its own executable path, not on PATH, so unpacking them apart from each other breaks job execution.

Put that directory on your PATH, or reference rewriter-queue’s full path directly (e.g. in Zed’s context_servers config or the Claude Code plugin’s MCP config).

No Windows build yet — job-queue’s worker uses Unix process-group APIs (libc::killpg) to reliably kill a job’s whole process tree on cancel or timeout, which doesn’t have a direct Windows equivalent as written. Runs fine under WSL in the meantime.

Building from source

Needs a recent stable Rust toolchain (rustup.rs).

git clone https://github.com/AndyGauge/rewriter-queue
cd rewriter-queue
cargo build --release

Binaries land in target/release/rewriter-queue and target/release/ai-org-orchestrator — same “keep them together” rule applies.

The lint plugins (optional)

ai-org-orchestrator’s quality gate can run ast-grep rules from ai-org-orchestrator/lints/rules/ as a non-blocking check alongside cargo build/clippy/test. It degrades gracefully if the binary isn’t installed — this is opt-in, not a hard dependency:

cargo install ast-grep --locked

The Zed extension and Claude Code plugin

Not part of the release archive — see Zed setup and claude-code-plugin/README.md for how to install each.

Zed setup

Two independent pieces: the context server (lets Zed’s agent panel submit and watch queue jobs) and the model provider (lets Zed chat with whatever backend is serving your models). You can set up either without the other.

1. The context server (zed-extension/)

This is a real Zed extension — zed-extension/extension.toml + zed-extension/src/lib.rs, compiled to WASM — not a hand-edited settings.json block. It registers rewriter-queue mcp as a context server and defaults its arguments to ["mcp"], so you don’t have to remember to add that yourself.

Build and install it as a dev extension:

rustup target add wasm32-wasip2
cargo build -p job-queue --release   # the rewriter-queue binary itself

In Zed: open the command palette and run zed: install dev extension, then point it at the zed-extension/ directory. Zed builds and loads the WASM extension itself — you don’t need to invoke the WASM build by hand.

Tell it where the rewriter-queue binary is, if it’s not on your PATH, in settings.json:

{
  "context_servers": {
    "rewriter-queue": {
      "command": {
        "path": "/absolute/path/to/rewriter-queue/target/release/rewriter-queue"
      }
    }
  }
}

Leave out command entirely if rewriter-queue is already on PATH (e.g. cargo install --path job-queue).

Point it at a queue server via ~/.config/rewriter/config.toml:

[queue]
url = "http://192.0.2.10:8003"   # wherever `rewriter-queue serve` is running

…and the token it expects, in ~/.config/rewriter/queue-token. Without a [queue] section, every command (including the ones the context server runs on your behalf) falls back to a local, file-backed queue — fine for trying this out on one machine, but you’ll usually want a real server if you’re also running the inference backend on a GPU box (see remote-deployment.md).

Once it’s connected (check the context-servers panel for a green status, not a red error dot), just ask Zed’s agent about your jobs in plain language — “what jobs are in the queue?”, “submit a synthesis run for ./src with a 20-iteration budget”, “download job 3’s artifacts to ~/Desktop” — and it calls the right tool (queue_submit/queue_list/queue_status/queue_artifacts/ queue_artifact/queue_download/queue_fetch/queue_cancel) itself. There’s no slash command for this — MCP tools are things the model calls on its own, not something you invoke directly, unless a server also exposes MCP prompts (this one doesn’t).

A live terminal view is also available without going through Zed’s agent at all: .zed/tasks.json in this repo defines a “Queue: watch” task that runs rewriter-queue watch.

2. The model provider

This is unrelated to the extension above — it’s Zed’s own openai_compatible provider type, pointed at whatever OpenAI-compatible server (llama.cpp, vLLM, etc.) you’re running your models on. Add it to settings.json:

{
  "language_models": {
    "openai_compatible": {
      "my-backend": {
        "api_url": "http://192.0.2.10:8001/v1",
        "available_models": [
          {
            "name": "your-model-id",
            "display_name": "My Model",
            "max_tokens": 128000,
            "capabilities": {
              "tools": true,
              "images": false,
              "parallel_tool_calls": false,
              "prompt_cache_key": false,
              "chat_completions": true
            }
          }
        ]
      }
    }
  }
}

"tools": true matters if you want Zed’s agent to actually call tools through this model — verified working with both llama.cpp’s --jinja mode and vLLM’s --enable-auto-tool-choice, but not every locally-served model reliably emits well-formed tool calls. If the agent starts guessing at shell commands instead of calling a tool by name, that’s usually the model’s tool-calling support, not a Zed or context-server bug — sanity-check by sending the same tool schema directly to the model’s /v1/chat/completions endpoint and see what comes back before assuming the wiring is broken.

Running two backends side by side (e.g. a general model and a coding-specialized one) on the same box just means two openai_compatible entries pointed at two different ports — see remote-deployment.md for sizing multiple backends into one machine’s memory.

Running this on a remote target

The queue server and the inference backend don’t have to be the same machine, and usually shouldn’t be — the inference backend wants a GPU, the queue server doesn’t need one at all. This covers running both on a remote Linux box (a rented GPU instance, a home server, whatever) and reaching it from your laptop.

The inference backend

Anything that speaks the OpenAI-compatible /v1/chat/completions API works — llama.cpp’s llama-server, vLLM, or a hosted API you don’t run yourself. If you’re self-hosting on a GPU box:

llama-server \
  -m /path/to/model.gguf \
  -a your-model-id \
  -ngl 999 \
  -c 65536 \
  -fa auto \
  --jinja \
  --host 0.0.0.0 --port 8001

Running more than one model on the same box is fine as long as the memory budget actually adds up — size it explicitly rather than finding out at OOM time. Two independent servers on two ports (e.g. 8001/8002) is simpler and safer than trying to multiplex one server across models. Watch in particular:

  • KV cache scales with context, not just model size. For a dense transformer: 2 (K and V) × layers × kv_heads × head_dim × context × bytes_per_element. A model with a Mixture-of-Experts or hybrid linear-attention architecture can have a very different (often much smaller) real footprint than its parameter count suggests — check the model’s own config.json, don’t assume.
  • -ctk q8_0 -ctv q8_0 (with -fa on) roughly halves KV cache size at a near-lossless quality cost — the quantization error is small and fixed per cached token, not cumulative with age, so it mostly costs you long-range recall precision in very long contexts, not general quality. q4_0 is meaningfully lossier; reach for q8_0 first.
  • Leave real headroom, not just enough to fit at idle — concurrent requests, the OS, and anything else on the box all need memory too.

The queue server

Runs anywhere that can reach the inference backend(s) over HTTP — doesn’t need to be the GPU box, and arguably shouldn’t be if you want to restart the inference backend independently of the queue.

rewriter-queue serve --bind 0.0.0.0:8003

It needs a bearer token (~/.config/rewriter/queue-token or REWRITER_QUEUE_TOKEN) and providers configured in ~/.config/rewriter/config.toml:

[[providers]]
name = "backend-a"
type = "openai-compat"
base_url = "http://<inference-box>:8001/v1"
models = ["your-model-id"]

[queue]
url = "http://<this-box>:8003"

A PATH gotcha that will bite you

quality_gate() shells out to cargo build/clippy/test (and, if installed, ast-grep) from inside the orchestrator subprocess the queue worker spawns. That subprocess inherits whatever PATH the queue server’s own process happened to have — which depends entirely on how you started it. A plain ssh host 'rewriter-queue serve ...' gets a minimal, non-login shell PATH on most systems (skips .bashrc/.profile, so ~/.cargo/bin usually isn’t on it) even though an interactive SSH session would have it. If cargo isn’t reliably found, every quality-gate check silently fails with “cargo build not available” — treated as a normal retryable error, so it doesn’t crash anything, it just quietly never actually validates code until the iteration budget runs out and broken output gets accepted anyway.

job-queue/src/worker.rs already guards against this — it explicitly sets PATH (prepending $HOME/.cargo/bin) on the orchestrator process it spawns, rather than trusting whatever it inherited. If you install cargo somewhere else, or need other tools (like ast-grep) on a non-standard PATH, that’s the place to adjust it.

Keeping it running

Neither binary is a managed service by default — write a systemd unit if you want one to survive a reboot or restart on crash. Two examples:

# /etc/systemd/system/rewriter-queue.service
[Unit]
Description=rewriter-queue server
After=network.target

[Service]
Type=simple
User=youruser
WorkingDirectory=/home/youruser/rewriter-queue
ExecStart=/home/youruser/rewriter-queue/target/release/rewriter-queue serve --bind 0.0.0.0:8003
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target
# /etc/systemd/system/llama-server.service
[Unit]
Description=llama.cpp inference server
After=network.target

[Service]
Type=simple
User=youruser
ExecStart=/home/youruser/llama.cpp/build/bin/llama-server \
  -m /home/youruser/models/your-model.gguf -a your-model-id \
  -ngl 999 -c 65536 -fa auto --jinja --host 0.0.0.0 --port 8001
Restart=on-failure
RestartSec=10

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now rewriter-queue llama-server

If you’d rather not deal with sudo/systemd at all, tmux (or screen) is a fine substitute for a personal box: start each in its own session so an SSH disconnect doesn’t kill it, and accept that a reboot needs a manual restart.

The lint plugins need ast-grep

ai-org-orchestrator/lints/ (see the root README) is optional — the orchestrator degrades gracefully if the binary isn’t found — but if you want it: cargo install ast-grep --locked on whichever machine actually runs the orchestrator subprocess (the queue server’s box, per the PATH note above).

Security posture — decide this deliberately

None of the above puts a token or TLS in front of the inference backend itself, and the queue server’s Bearer token is the only thing gating job submission (no per-job auth, no rate limiting). That’s a reasonable default for a private LAN or a locked-down VPC where you trust everything that can reach the port — it is not a default you want facing the public internet. If you’re putting either behind a real network boundary (a cloud security group, a home router’s port forward, anything multi-tenant), add a reverse proxy with TLS and real auth in front, and give llama-server its own --api-key rather than relying on network trust alone.