rewriter-queue
A resumable job queue for multi-agent LLM code synthesis. You submit a source tree and a mission; a small organization of agent roles (implementer, security/maintenance/production-readiness reviewers, a merge agent for large inputs) iterates on it — patching rather than rewriting where it can, checkpointing every real step so an interrupted run resumes instead of starting over — until the reviewers are satisfied or the iteration budget runs out.
Submit and watch jobs from a terminal, from Zed, or from Claude Code.
Not a hosted service — three small Rust binaries you run yourself, against whatever OpenAI-compatible inference endpoint you point them at.
Start with the tutorial.
Tutorial
This walks through the whole loop once, end to end: get a binary, point it at a model, submit something to rewrite, and watch it work. Fifteen minutes if you’ve already got an OpenAI-compatible inference server running somewhere; longer if you need to stand one up first (see Remote deployment for that part).
1. Get the binaries
Download the archive for your platform from the
latest release
and unpack it — see Installation if you’d rather build
from source. Either way you end up with two binaries sitting next to each
other: rewriter-queue and ai-org-orchestrator. They need to stay in the
same directory — the queue’s worker finds the orchestrator by looking beside
its own binary, not on PATH.
Put that directory on your PATH, or just remember the full path for the
commands below.
2. Tell it where your model lives
Create ~/.config/rewriter/config.toml:
[[providers]]
name = "local"
type = "openai-compat"
base_url = "http://localhost:8001/v1"
models = ["your-model-id"]
Point base_url at any OpenAI-compatible /v1/chat/completions endpoint —
a llama.cpp/vLLM server you’re running, or a hosted API. If you don’t
have one yet, that’s the whole subject of
Remote deployment.
3. Pick a queue mode
For trying this out on one machine, skip the server entirely — every
command below falls back to a local, file-backed queue when there’s no
[queue] section in the config. If you’re running the inference backend on
a separate box and want to submit jobs from your laptop, start the server
there instead and add a [queue] url = "..." pointing at it (see the root
README’s quickstart for the server command and token setup).
This tutorial uses local mode — nothing extra to start.
4. Submit something
Pick a small, real piece of code you want rewritten or reviewed — a single source file or a small directory is a better first run than a whole project. The workspace directory holds the tool’s working state as it goes, plus one thing you provide before submitting: the mission.
The mission is a required artifact — a short statement of what V2 should
actually do, not just “port this file.” Without it, the first agent in the
pipeline has to guess your intent from raw source alone, which shows up
later as a contract that’s technically correct but not what you meant. It’s
stored the same way as every other artifact this pipeline produces
(objective_contract.md, schema.rs, …), at <workspace>/artifacts/mission.md,
and the run refuses to start without one:
mkdir -p /tmp/my-first-run/artifacts
cat > /tmp/my-first-run/artifacts/mission.md <<'EOF'
# Mission
V1 is a small in-memory counter store. Rebuild it as V2 with identical
behavior: increment a key's count, and read a key's current count (0 for
an unseen key).
EOF
rewriter-queue submit \
--source /path/to/the/code \
--workspace /tmp/my-first-run \
--max-iter 10
That prints something like queued job #1 (0 ahead). In local mode with
nothing else queued, it starts immediately.
5. Watch it work
rewriter-queue watch
A live view that refreshes once a second: which job is running, what stage
it’s on (ObjectiveContract → test matrix → inductive analysis → schema →
V2 synthesis), and how long it’s been going. Ctrl-C to stop watching —
the job keeps running either way.
Want more detail on one job specifically?
rewriter-queue status 1
Shows the last 20 lines of its log, including which agent is running and the token counts for each call.
6. See what it’s produced, before it’s even done
You don’t have to wait for the job to finish to look at what it’s built so far:
rewriter-queue artifacts 1 # list what's been emitted
rewriter-queue artifact 1 objective_contract.md # read one of them
objective_contract.md is usually the first thing worth reading — it’s the
agreed-on shape of what’s being built (public API, invariants, edge cases,
explicit exclusions) before any code gets written.
7. Get the result
In local mode there’s nothing to download — the workspace directory you
passed to submit is where everything lands, no copy needed:
ls /tmp/my-first-run/artifacts/v2/ # the actual synthesized code
(rewriter-queue download 1 --out ./result is for the remote-server case —
see Remote deployment — where the job ran on a
different machine and you need its artifacts/ tree copied to yours in one
call instead of one file at a time.)
8. If it gets interrupted
Kill rewriter-queue serve (or just Ctrl-C a local-mode run) mid-job and
resubmit isn’t what you want — restart the server (or just try watch
again in local mode) and the same job picks back up from wherever it had
already checkpointed, not from the beginning. This is true down to the
level of individual agent calls, not just whole pipeline stages — see the
root README’s design notes if you’re curious how.
What’s next
- Installation — the binary download details, and building from source if you’d rather.
- Zed setup — submit and watch jobs from Zed’s agent panel instead of a terminal.
- The Claude Code plugin — the same queue tools inside Claude Code, plus each reviewer persona as a standalone slash command.
- Remote deployment — running the inference backend and the queue server on a separate (e.g. GPU) machine from the one you’re submitting jobs from.
Installation
Prebuilt binaries
Every release ships prebuilt archives for:
| Platform | Archive |
|---|---|
| Linux x86_64 | rewriter-queue-x86_64-unknown-linux-gnu.tar.gz |
| Linux aarch64 (ARM64) | rewriter-queue-aarch64-unknown-linux-gnu.tar.gz |
| macOS Intel | rewriter-queue-x86_64-apple-darwin.tar.gz |
| macOS Apple Silicon | rewriter-queue-aarch64-apple-darwin.tar.gz |
curl -LO https://github.com/AndyGauge/rewriter-queue/releases/latest/download/rewriter-queue-<your-target>.tar.gz
tar xzf rewriter-queue-<your-target>.tar.gz
Each archive contains two binaries — rewriter-queue and
ai-org-orchestrator — plus the license files. Keep both binaries in the
same directory. The queue’s worker finds the orchestrator by looking
beside its own executable path, not on PATH, so unpacking them apart from
each other breaks job execution.
Put that directory on your PATH, or reference rewriter-queue’s full
path directly (e.g. in Zed’s context_servers config or the Claude Code
plugin’s MCP config).
No Windows build yet — job-queue’s worker uses Unix process-group APIs
(libc::killpg) to reliably kill a job’s whole process tree on cancel or
timeout, which doesn’t have a direct Windows equivalent as written. Runs
fine under WSL in the meantime.
Building from source
Needs a recent stable Rust toolchain (rustup.rs).
git clone https://github.com/AndyGauge/rewriter-queue
cd rewriter-queue
cargo build --release
Binaries land in target/release/rewriter-queue and
target/release/ai-org-orchestrator — same “keep them together” rule
applies.
The lint plugins (optional)
ai-org-orchestrator’s quality gate can run ast-grep
rules from ai-org-orchestrator/lints/rules/ as a non-blocking check
alongside cargo build/clippy/test. It degrades gracefully if the
binary isn’t installed — this is opt-in, not a hard dependency:
cargo install ast-grep --locked
The Zed extension and Claude Code plugin
Not part of the release archive — see Zed setup and
claude-code-plugin/README.md
for how to install each.
Zed setup
Two independent pieces: the context server (lets Zed’s agent panel submit and watch queue jobs) and the model provider (lets Zed chat with whatever backend is serving your models). You can set up either without the other.
1. The context server (zed-extension/)
This is a real Zed extension — zed-extension/extension.toml +
zed-extension/src/lib.rs, compiled to WASM — not a hand-edited
settings.json block. It registers rewriter-queue mcp as a context
server and defaults its arguments to ["mcp"], so you don’t have to
remember to add that yourself.
Build and install it as a dev extension:
rustup target add wasm32-wasip2
cargo build -p job-queue --release # the rewriter-queue binary itself
In Zed: open the command palette and run zed: install dev extension,
then point it at the zed-extension/ directory. Zed builds and loads the
WASM extension itself — you don’t need to invoke the WASM build by hand.
Tell it where the rewriter-queue binary is, if it’s not on your
PATH, in settings.json:
{
"context_servers": {
"rewriter-queue": {
"command": {
"path": "/absolute/path/to/rewriter-queue/target/release/rewriter-queue"
}
}
}
}
Leave out command entirely if rewriter-queue is already on PATH
(e.g. cargo install --path job-queue).
Point it at a queue server via ~/.config/rewriter/config.toml:
[queue]
url = "http://192.0.2.10:8003" # wherever `rewriter-queue serve` is running
…and the token it expects, in ~/.config/rewriter/queue-token. Without a
[queue] section, every command (including the ones the context server
runs on your behalf) falls back to a local, file-backed queue — fine for
trying this out on one machine, but you’ll usually want a real server if
you’re also running the inference backend on a GPU box (see
remote-deployment.md).
Once it’s connected (check the context-servers panel for a green status,
not a red error dot), just ask Zed’s agent about your jobs in plain
language — “what jobs are in the queue?”, “submit a synthesis run for
./src with a 20-iteration budget”, “download job 3’s artifacts to
~/Desktop” — and it calls the right tool
(queue_submit/queue_list/queue_status/queue_artifacts/
queue_artifact/queue_download/queue_fetch/queue_cancel) itself.
There’s no slash command for this — MCP tools are things the model calls on
its own, not something you invoke directly, unless a server also exposes
MCP prompts (this one doesn’t).
A live terminal view is also available without going through Zed’s agent
at all: .zed/tasks.json in this repo defines a “Queue: watch” task
that runs rewriter-queue watch.
2. The model provider
This is unrelated to the extension above — it’s Zed’s own
openai_compatible provider type, pointed at whatever OpenAI-compatible
server (llama.cpp, vLLM, etc.) you’re running your models on. Add it to
settings.json:
{
"language_models": {
"openai_compatible": {
"my-backend": {
"api_url": "http://192.0.2.10:8001/v1",
"available_models": [
{
"name": "your-model-id",
"display_name": "My Model",
"max_tokens": 128000,
"capabilities": {
"tools": true,
"images": false,
"parallel_tool_calls": false,
"prompt_cache_key": false,
"chat_completions": true
}
}
]
}
}
}
}
"tools": true matters if you want Zed’s agent to actually call tools
through this model — verified working with both llama.cpp’s --jinja
mode and vLLM’s --enable-auto-tool-choice, but not every locally-served
model reliably emits well-formed tool calls. If the agent starts guessing
at shell commands instead of calling a tool by name, that’s usually the
model’s tool-calling support, not a Zed or context-server bug — sanity-check
by sending the same tool schema directly to the model’s /v1/chat/completions
endpoint and see what comes back before assuming the wiring is broken.
Running two backends side by side (e.g. a general model and a
coding-specialized one) on the same box just means two openai_compatible
entries pointed at two different ports — see
remote-deployment.md for sizing multiple backends
into one machine’s memory.
Running this on a remote target
The queue server and the inference backend don’t have to be the same machine, and usually shouldn’t be — the inference backend wants a GPU, the queue server doesn’t need one at all. This covers running both on a remote Linux box (a rented GPU instance, a home server, whatever) and reaching it from your laptop.
The inference backend
Anything that speaks the OpenAI-compatible /v1/chat/completions API
works — llama.cpp’s llama-server, vLLM, or a hosted API you don’t run
yourself. If you’re self-hosting on a GPU box:
llama-server \
-m /path/to/model.gguf \
-a your-model-id \
-ngl 999 \
-c 65536 \
-fa auto \
--jinja \
--host 0.0.0.0 --port 8001
Running more than one model on the same box is fine as long as the
memory budget actually adds up — size it explicitly rather than finding out
at OOM time. Two independent servers on two ports (e.g. 8001/8002) is
simpler and safer than trying to multiplex one server across models. Watch
in particular:
- KV cache scales with context, not just model size. For a dense
transformer:
2 (K and V) × layers × kv_heads × head_dim × context × bytes_per_element. A model with a Mixture-of-Experts or hybrid linear-attention architecture can have a very different (often much smaller) real footprint than its parameter count suggests — check the model’s ownconfig.json, don’t assume. -ctk q8_0 -ctv q8_0(with-fa on) roughly halves KV cache size at a near-lossless quality cost — the quantization error is small and fixed per cached token, not cumulative with age, so it mostly costs you long-range recall precision in very long contexts, not general quality.q4_0is meaningfully lossier; reach forq8_0first.- Leave real headroom, not just enough to fit at idle — concurrent requests, the OS, and anything else on the box all need memory too.
The queue server
Runs anywhere that can reach the inference backend(s) over HTTP — doesn’t need to be the GPU box, and arguably shouldn’t be if you want to restart the inference backend independently of the queue.
rewriter-queue serve --bind 0.0.0.0:8003
It needs a bearer token (~/.config/rewriter/queue-token or
REWRITER_QUEUE_TOKEN) and providers configured in
~/.config/rewriter/config.toml:
[[providers]]
name = "backend-a"
type = "openai-compat"
base_url = "http://<inference-box>:8001/v1"
models = ["your-model-id"]
[queue]
url = "http://<this-box>:8003"
A PATH gotcha that will bite you
quality_gate() shells out to cargo build/clippy/test (and, if
installed, ast-grep) from inside the orchestrator subprocess the queue
worker spawns. That subprocess inherits whatever PATH the queue server’s
own process happened to have — which depends entirely on how you started
it. A plain ssh host 'rewriter-queue serve ...' gets a minimal, non-login
shell PATH on most systems (skips .bashrc/.profile, so ~/.cargo/bin
usually isn’t on it) even though an interactive SSH session would have it.
If cargo isn’t reliably found, every quality-gate check silently fails with
“cargo build not available” — treated as a normal retryable error, so it
doesn’t crash anything, it just quietly never actually validates code
until the iteration budget runs out and broken output gets accepted anyway.
job-queue/src/worker.rs already guards against this — it explicitly sets
PATH (prepending $HOME/.cargo/bin) on the orchestrator process it
spawns, rather than trusting whatever it inherited. If you install cargo
somewhere else, or need other tools (like ast-grep) on a non-standard
PATH, that’s the place to adjust it.
Keeping it running
Neither binary is a managed service by default — write a systemd unit if you want one to survive a reboot or restart on crash. Two examples:
# /etc/systemd/system/rewriter-queue.service
[Unit]
Description=rewriter-queue server
After=network.target
[Service]
Type=simple
User=youruser
WorkingDirectory=/home/youruser/rewriter-queue
ExecStart=/home/youruser/rewriter-queue/target/release/rewriter-queue serve --bind 0.0.0.0:8003
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
# /etc/systemd/system/llama-server.service
[Unit]
Description=llama.cpp inference server
After=network.target
[Service]
Type=simple
User=youruser
ExecStart=/home/youruser/llama.cpp/build/bin/llama-server \
-m /home/youruser/models/your-model.gguf -a your-model-id \
-ngl 999 -c 65536 -fa auto --jinja --host 0.0.0.0 --port 8001
Restart=on-failure
RestartSec=10
[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable --now rewriter-queue llama-server
If you’d rather not deal with sudo/systemd at all, tmux (or screen)
is a fine substitute for a personal box: start each in its own session so
an SSH disconnect doesn’t kill it, and accept that a reboot needs a manual
restart.
The lint plugins need ast-grep
ai-org-orchestrator/lints/ (see the root README) is optional — the
orchestrator degrades gracefully if the binary isn’t found — but if you
want it: cargo install ast-grep --locked on whichever machine actually
runs the orchestrator subprocess (the queue server’s box, per the PATH
note above).
Security posture — decide this deliberately
None of the above puts a token or TLS in front of the inference backend
itself, and the queue server’s Bearer token is the only thing gating job
submission (no per-job auth, no rate limiting). That’s a reasonable
default for a private LAN or a locked-down VPC where you trust everything
that can reach the port — it is not a default you want facing the public
internet. If you’re putting either behind a real network boundary (a cloud
security group, a home router’s port forward, anything multi-tenant), add
a reverse proxy with TLS and real auth in front, and give llama-server
its own --api-key rather than relying on network trust alone.