Multi-turn long-horizon planning is a crucial aspect of foundation model agents, but its development remains poorly understood due to the opaque nature of internet-based training data. Researchers have introduced a novel, unified environment that facilitates precise control over multi-turn interactions, enabling a deeper understanding of how planning abilities are acquired and integrated1. This controlled environment allows for the examination of single- and multi-teacher on-policy agentic distillation, a process that distills knowledge from pre-trained models to improve planning capabilities. By leveraging this framework, scientists can investigate the physics of long-horizon planning, from pre-training to post-training, and gain insights into the underlying mechanisms that drive planning ability. This matters to practitioners because it has the potential to significantly improve the planning capabilities of foundation model agents, leading to more effective and efficient decision-making in complex, real-world scenarios.