Scored
bets.
Dated predictions scored in public. I don't claim judgment — I show how the bets graded. Misses included.
4 bets · 1 scored · 3 open
- BET-01
2026.07A single integration-directed cue rescues parallel agent assembly failures more than behavioral contracts — on frontier models in a step-by-step build harness.
Confirmed on Claude Opus/Sonnet in agentic harness (21% → 82% on one seam). Fails in single-shot generation and does not transfer cleanly across model families without the harness — so the bet holds with explicit scope, not as a universal law.
researchpartial · 2026.07 - BET-02
2026.07Agents can enumerate their own integration failures on request reliably enough to compile that enumeration directly into an executable gate — skipping the fragile prompt step.
researchopen - BET-03
2026.07A cue aimed at a specific bad habit (e.g. caching when freshness is required) beats a general integration review cue on the stubborn failure modes.
researchopen - BET-04
2026.07The integration-cue effect replicates on a small Claude model that could not be run in the first experiment round.
researchopen