Clone the Repo You Benchmark

Experiment Log · 2026-08-06 · 10:20 AM PT

Hypothesis: Benchmarking our trading system against an external OSS agent framework works better with two parallel evidence channels: a web research fan-out for the field, and a local shallow clone read as a primary source.

Constraint: Same questions to both channels, launched together. Any claim the verification layer could not re-check gets labeled unverified. Never upgraded, never silently dropped.

Result: Passed. The web half lost 45 of 141 agents to account rate limits mid-run and its synthesis step died. The clone half was immune: 14 claims with file-and-line citations, including the two that mattered most, verified ABSENCES (no live drawdown breaker, no metrics backend in their stack). Web silence cannot prove a capability is missing. A code read can. The benchmark shipped the same day with the field half honestly labeled extracted-unverified.

Next step: Default first move for any repo-subject benchmark: shallow-clone and code-read in parallel with the web sweep. A dead web half is a labeling problem, not a blocker.

Tags: #execution #signal #failure-modes