Researchers have developed ReflectRL, a novel approach to on-policy training that leverages golden negative trajectories to enhance the reasoning capabilities of large language models. By incorporating reflective-to-direct reasoning, ReflectRL enables models to learn from failed expert trajectories, which are typically discarded as negative examples. This approach allows models to improve their performance on harder problems, where expert models often fail. ReflectRL builds upon existing trajectory-guided methods, which rely on supervision from stronger expert models, but are limited by the lack of guidance when the expert fails1. By learning from golden negative trajectories, ReflectRL can improve the robustness and generalizability of large language models. This advancement has significant implications for the development of more sophisticated AI systems, which can impact various aspects of society, including policy, security, and workforce dynamics. The ability to learn from failure can lead to more resilient and adaptable AI models, making ReflectRL a crucial step forward in AI research.