Exploring Next / Topics / Difficulty Aware Compute Allocation Topic Difficulty Aware Compute Allocation 1 episode Ep 986 Sep 18, 2026 Learning Difficulty Aware Length Controlfor Efficient Hybrid Reasoning Models Edmund and Geffen dig into When2Think, a framework that teaches large reasoning models to spend fewer tokens on easy math problems while still thinking deeply on hard ones, using a clever difficulty-aware reward signal instead of a separate controller or reward model. InferenceTrainingChain Of ThoughtReinforcement Learning From Human Feedback