Researchers have introduced Output Reset (OR), a novel differentiable trust region approach for policy optimization, as a potential alternative to clipped surrogate objectives used in Proximal Policy Optimization (PPO) and other reinforcement learning algorithms. By replacing the clipped policy term with an OR squared-margin loss, PPO-OR and GRPO-OR aim to mitigate the abrupt change in the scalar objective's derivative caused by favorable-direction saturation. This smooth one-sided saturation rule may offer improved performance in large language model post-training1. The OR method allows for a more nuanced control over the policy updates, potentially leading to more stable and efficient learning. As AI models become increasingly prevalent, advancements in reinforcement learning have significant implications for various domains, including security and workforce dynamics. The development of OR highlights the ongoing efforts to improve the robustness and reliability of AI systems, which is crucial for ensuring their safe and effective deployment.