Efficient long-video understanding hinges on the ability of vision-language models to selectively focus on sparse visual evidence. Existing methods rely on static frame selection, which can be limiting, or agent-based schedulers that incur significant computational costs. EcoFrame, a novel approach, aims to strike a balance between efficiency and adaptivity by dynamically scheduling visual evidence. This adaptive visual evidence scheduling enables models to reason over a small, strategically selected subset of frames, rather than relying on fixed budgets or exhaustive search. By doing so, EcoFrame achieves a more efficient and effective understanding of long videos1. The implications of this research extend beyond the realm of computer vision, as it can inform the development of more sophisticated threat detection systems. So what matters to practitioners is that EcoFrame's adaptive scheduling can be leveraged to enhance the efficiency and accuracy of video analysis in various applications, including security and surveillance.