Multimodal large language models (MLLMs) face a significant challenge known as Modality Imbalance, where textual information overshadows other input sources, hindering visual reasoning capabilities. To address this, researchers have proposed OPD-V, a visual on-policy self-distillation method that incorporates modality balance. This approach enables MLLMs to effectively leverage diverse input sources, mitigating the dominance of textual information and enhancing overall visual reasoning. By doing so, OPD-V improves the model's ability to generate more balanced and informative outputs. The introduction of modality balance in OPD-V has the potential to significantly impact the performance of MLLMs in various applications, including those related to policy, security, and workforce dynamics1. This development matters to practitioners as it highlights the need to consider modality balance when designing and training MLLMs to ensure more accurate and reliable outputs.