Topic

Gpt 5

4 episodes

  1. Ep 668

    Stripe Benchmark Shows AI Agents Build Integrations but Struggle with Validation

    Cathy is skeptical that the Stripe benchmark proves much beyond a familiar split: agents can write integration code, but they still get tripped up by validation, state, and recovery. Jessica thinks that’s exactly the useful part, because in real product work the hard failure is often whether the thing can prove it worked, not whether it can type out the API calls.

  2. Ep 367

    From Context to Skills: Can Language Models Learn from Context Skillfully?

    Cody and Justy dig into Ctx2Skill, a self-evolving framework that turns long, dense context into reusable natural-language skills for language models. They talk through the core loop, the role of Challenger, Reasoner, Judge, and the replay trick that keeps the system from drifting into weird overfit territory, then land on what it means for product teams trying to ship context-heavy workflows.

  3. Ep 169

    Thinking in Frames: How Visual Context and Test Time Scaling Empower Video Reasoning

    Today, we dive into a game-changing approach to visual reasoning in video generation. How does this solve real-world problems?

  4. Ep 15

    GPT 5 prompting guide | OpenAI Cookbook

    Unlock the full potential of GPT-5 with practical prompting strategies to enhance performance and steerability.