Researchers have made a breakthrough in reinforcement learning with verifiable rewards, a technique used to improve reasoning in large language models. Typically, these models struggle with difficult problems because they receive no learning signal when they cannot generate correct solutions. To overcome this, scientists have introduced a new approach called Off-Context GRPO, which provides privileged guidance during training, such as solution prefixes, to steer the model towards correct solutions1. This method enables the model to learn from hard problems, even when it cannot generate any correct solutions on its own. The development of more advanced language models using reinforcement learning has significant implications for security, as it can both enhance capabilities and increase risk surfaces. As large language models become more powerful, their potential impact on security landscapes grows, making it essential for practitioners to stay informed about the latest advancements and their potential consequences.