rust-skeptic 0.14: Six Systems Problems Hiding in a Doc Tester
TL;DR — skeptic 0.14.0 is out. It turns the Rust code blocks in your markdown into tests, and this release is mostly the work from my last post plus what code review and a real-world run turned up afterwards. Cold-start setup is up to about 2x faster than the branch just before the review fixes, the Rust Cookbook’s 225 snippets pass, and one quadratic-time bug (1.1 s for a 1000-snippet page, now milliseconds), and a few correctness bugs that only show up with a messy target/ are fixed. This post sorts the changes by the kind of systems problem each one is, because that turned out to be the most useful way to think about them.
What’s in 0.14
- Finds dependencies on Cargo 1.77 and later, and supports the newer
build/<pkg>/<hash>/directory layout. - Test extraction is linear in page size instead of quadratic.
- Resolves rustc setup once per process and runs snippets on a bounded worker pool (
SKEPTIC_JOBS). - Cache file is JSON instead of the unmaintained
bincode; old caches are rebuilt. - Dependency cache is validated before reuse and written atomically.
- Picks the right build of a crate when several of the same version exist.
- A panicking snippet no longer hangs the rest of the suite.
- Minimum Rust version is 1.85.
once_cell,num_cpusandsemverare gone in favour ofstd.
How I measured
The benchmarks live in a separate crate using divan through CodSpeed’s codspeed-divan-compat, so the same code runs under cargo bench and in CI. Two targets:
generateis in-process: parse markdown, emit tests. CodSpeed runs this in CI, and its history is what the first section below comes from.rtis the per-snippet phase. It spawnscargo metadataandrustc, so I run it locally against a built cookbook (COOKBOOK_DIR).
The numbers below are from one machine, single runs, comparing the commit I measured last time against 0.14. Medians. Treat differences of a few percent as noise.

1. Algorithmic complexity: the one CodSpeed caught
The benchmark that mattered most wasn’t the one I expected. PR #4 added CodSpeed benchmarks to the unmodified code, and one of them feeds a single markdown page with a growing number of code blocks to generate_doc_tests. Going from 10 blocks to 100 to 1000 should cost about 10x each time. It didn’t:
| Blocks in one page | Unmodified (PR #4) | After the threading work | 0.14 |
|---|---|---|---|
| 10 | 479.7 µs | 206.3 µs | 107.3 µs |
| 100 | 9 ms | 1.5 ms | 319.8 µs |
| 1000 | 1.1 s | 13.9 ms | 2.4 ms |
Going from 100 to 1000 blocks cost 122x more on the unmodified code, and about 9x more afterwards. The cause was in the generator, not the runtime: line numbers were computed by counting newlines from the start of the file for every markdown event, so extracting tests was quadratic in page size. Only code blocks need line numbers, so it now counts from the previous code block instead (the commit is eleven lines). A real cookbook page has a handful of snippets, so you’d never notice it there, but a long reference page would have paid for it on every build.
Caveats on the table. The first two columns are from CodSpeed’s instruction-counting simulation. The last is from the Walltime instrument on a shared hosted runner, which is a different measurement and noisier, so the step from the middle column to the last isn’t a like-for-like comparison. The step from the first to the middle is, and it’s the big one. The whole-cookbook benchmark went from 11.4 ms to 5 ms across the same change of instrument, and has no baseline in PR #4 (its set of benchmarks was different).
2. Redundant work: do it once
cargo metadata is slow, and the setup path ran it three or more times, under a global lock that every test thread waited on. Now it runs once per setup and the result is passed down. Package lookups are indexed instead of scanned.
| Cold start (first snippet) | Before | 0.14 |
|---|---|---|
| no dependencies | 123 ms | 83 ms |
| 12 dependencies | 420 ms | 238 ms |
two versions of rand |
200 ms | 106 ms |
| 8-member workspace | 351 ms | 146 ms |
Warm per-snippet numbers didn’t move (compile_test 77 → 76 ms, run_test 466 → 459 ms), which is the right result: once setup is cached, what remains is rustc. This is a fixed cost paid once per test process, so it matters most for large dependency trees and for anyone running tests repeatedly.
3. Cache coherence: a cache is a claim about the past
The disk cache was keyed on paths and the mtimes of Cargo.toml and Cargo.lock, and trusted for an hour. That key says nothing about the thing the cache actually points at. If a dependency is rebuilt with different features or flags, or cargo clean runs, the cached libfoo-<oldhash>.rlib is gone, and every snippet fails until the entry ages out.
Two fixes. A cached entry is only reused if every rlib it names still exists. And the file is written to a temporary file and renamed, so two test processes can’t leave a half-written cache behind. Neither changes a benchmark, because they’re about being correct when the cache is wrong. A speedup that sometimes serves stale answers isn’t a speedup.
4. Identity: which serde is the right serde?
This was the most interesting bug, and I found it by running the real cookbook against the release candidate.
A target/ directory can hold several builds of the same crate at the same version, differing in features or in whether they were built for build scripts. rustc only accepts the build the other crates were compiled against. If a snippet uses csv and serde together and the two disagree, you get an error like Record: serde::Deserialize<'de> is not satisfied, which looks like a bug in your code and isn’t.
My first fix, “prefer the freshest build”, failed 13 of the cookbook’s 225 snippets. The branch before it, which took whichever path sorted last, passed, by luck. Neither rule looks at the thing that actually matters.
Cargo does record it. Each unit’s fingerprint lists the builds of its dependencies it used. So skeptic now counts, for each candidate build, how many other units were compiled against it (including the project’s own test targets), and links the most depended-upon one, with freshness only breaking ties. Telling apart two versions of the same crate now reads the unit’s dep-info file, which lists the source files it was built from, instead of guessing from feature names. The rand-specific special case is gone.
With that, the cookbook’s 225 snippets pass again; a full cargo test --test skeptic takes about 57 s here, against 297 s for the fork the cookbook originally pinned.
5. Bounded concurrency and fault isolation
libtest gives every test its own thread, so without a limit a large suite launches that many rustc processes at once. The worker pool from the last post bounds it. The fix this time was the failure mode: a worker that panicked (a non-UTF-8 byte in output, a failed spawn) died quietly, the pool was never refilled, and the remaining tests waited forever. Workers now catch the panic and report a failure for that one snippet. Setup errors become test failures rather than panics.

6. A measurement problem: what the benchmark can’t see
CodSpeed’s default mode counts instructions. That’s stable, but it can’t see time spent in system calls. After I added the benchmarks, CodSpeed flagged three of them as spending significant time in syscalls, which means the numbers understated their real cost. The fix is its Walltime instrument, which measures elapsed time.
The catch is that walltime on a shared hosted CI runner is noisy, and CodSpeed says so. For the markdown-generation benchmarks it reports about 5 ms for the whole cookbook and 2.4 ms for a 1000-snippet page, but I wouldn’t read small differences into numbers from that environment. Stable walltime numbers need a dedicated macro runner. The rt benchmark spawns cargo and rustc, which instruction counting can’t see, and it needs a built cookbook, so I still run it locally. Choosing the instrument is part of the result: the wrong one gives you a precise, repeatable measurement of the wrong thing.
What I’d still like to fix
- Skip
cargo metadataentirely when the disk cache is valid, by storing the edition and a lock-file check with the entry. - Move
.skeptic-cacheout of the source tree and intotarget/, where it belongs. That’s a behaviour change, so it didn’t go into 0.14. - Batch snippets that share a template into one compile.
skeptic 0.14 is on crates.io, the code is in AndyGauge/rust-skeptic, and the changes landed through PR #3.