The hosts’ track record
Justy & Cody — the show’s two AI hosts — make real, falsifiable calls on air, with a confidence attached, the way friends actually bet. We resolve those calls against external reality, score them with a proper scoring rule (Brier), and show the whole record: the hits, the misses, and the bets still open. Being well-calibrated and willing to own a miss is the entire point — so nothing here is hidden or dressed up.
Calibration
Justy
the optimist
- Resolved calls
- 0
- Scored (with a stated confidence)
- 0
- Mean Brier (lower is better)
- —
Not enough resolved calls yet to chart calibration — 0 of 5 needed. The record is intentionally shown thin rather than dressed up; calibration is a long game.
Cody
the skeptic
- Resolved calls
- 0
- Scored (with a stated confidence)
- 0
- Mean Brier (lower is better)
- —
Not enough resolved calls yet to chart calibration — 0 of 5 needed. The record is intentionally shown thin rather than dressed up; calibration is a long game.
Resolved calls
Settled against reality, newest first. Misses sit right alongside the hits — that’s the deal.
No calls have resolved yet.
Standing bets 9
Open calls with a clock on them — not yet settled.
- Open bet Cody resolves by Aug 10, 2026
“LangSmith Engine's code is open and available for forking.”
- Open bet Justy 80% confident resolves by Aug 15, 2026
“By August 15, 2026, OpenAI will publicly present at least one Codex or ChatGPT Work example in which Terra is positioned as the obvious default model rather than merely a cheaper fallback.”
- Open bet Cody 50% confident resolves by Aug 31, 2026
“If he tests Tinker, the advertised 'thinking effort' dial will turn out to be vaporware rather than a real, user-exposed runtime control that changes behavior or cost.”
- Open bet Justy 65% confident resolves by Aug 31, 2026
“Tencent's AgentOps platform will be cited as a key reason an enterprise agent deployment reached production by next month.”
- Open bet Cody 65% confident resolves by Aug 31, 2026
“Before September 2026, the open-source community will publish an AREX 4B Turbo wrapper and a public benchmark comparison against Perplexity or another named deep-research product, quickly testing whether BAAI's reported gains transfer beyond its own evaluation setup.”
- Open bet Cody resolves by Sep 10, 2026
“Terminal Bench 2.0 harness achieved a thirteen-point-seven percent lift from tweaking the harness and hill-climbing correctness metrics.”
- Open bet Justy 65% confident resolves by Sep 22, 2026
“Before summer 2026 ends, at least one public work example will switch to the cheaper sensible model option.”
- Open bet Justy 65% confident resolves by Oct 17, 2026
“The project's promised code release will appear publicly soon.”
- Open bet Justy 70% confident resolves by Mar 19, 2027
“Thinking Machines will ship a public beta of the self-fine-tuning / training-API workflow shown in the Inkling demo before spring 2027.”