/lib/record

Scored
bets.

Dated predictions scored in public. I don't claim judgment — I show how the bets graded. Misses included.

4 bets · 1 scored · 3 open

  • BET-01
    2026.07

    A single integration-directed cue rescues parallel agent assembly failures more than behavioral contracts — on frontier models in a step-by-step build harness.

    Confirmed on Claude Opus/Sonnet in agentic harness (21% → 82% on one seam). Fails in single-shot generation and does not transfer cleanly across model families without the harness — so the bet holds with explicit scope, not as a universal law.

    researchpartial · 2026.07
  • BET-02
    2026.07

    Agents can enumerate their own integration failures on request reliably enough to compile that enumeration directly into an executable gate — skipping the fragile prompt step.

    researchopen
  • BET-03
    2026.07

    A cue aimed at a specific bad habit (e.g. caching when freshness is required) beats a general integration review cue on the stubborn failure modes.

    researchopen
  • BET-04
    2026.07

    The integration-cue effect replicates on a small Claude model that could not be run in the first experiment round.

    researchopen