The relentless march of artificial intelligence continues, this time making significant strides in video understanding. A new model called VideoPro, detailed in a recent paper on arXiv, is making waves by tackling the notoriously difficult challenge of long-form video question answering (videoQA). What sets VideoPro apart is its adaptive reasoning approach, cleverly balancing speed and accuracy to deliver state-of-the-art performance.
Fast and Slow Reasoning for Optimal Efficiency
VideoPro employs a "fast-slow" reasoning framework inspired by human cognition. For simple queries, the system utilizes efficient VideoLLMs for a quick answer. However, when faced with more complex questions, VideoPro intelligently invokes a more thorough visual program reasoning approach. "Simple queries are directly solved by VideoLLMs, while difficult ones invoke visual program reasoning, motivated by human-like reasoning processes," the researchers explain. This two-tiered approach optimizes for both speed and accuracy, ensuring that resources are allocated efficiently.
Importantly, the model also incorporates a fallback mechanism. If the slow-reasoning process encounters a snag, it can revert to the faster method, preventing complete stalls. The system refines its performance through parameter search during both training and inference. By tweaking the parameters of the visual modules, multiple program variants are generated. During training, the models that produce correct answers are favored. During inference, the model selects the variant that generates the most confident result.
Benchmarking and Open Source Alignment
To train VideoPro, the team developed a diverse and high-quality fast-slow reasoning dataset. This dataset is designed to align open-source language models with the ability to generate effective visual program workflows. The results are impressive. VideoPro achieves 50.4% accuracy on LVBench, surpassing GPT-4o and matching the performance of Qwen2.5VL-72B on VideoMME. These benchmarks demonstrate VideoPro's ability to compete with even the most advanced closed-source models.
The success of VideoPro highlights the potential of adaptive reasoning frameworks in AI. By mimicking human cognitive processes, these models can achieve greater efficiency and accuracy in complex tasks. As AI continues to evolve, we can expect to see more innovations that draw inspiration from the way humans think and learn. Moreover, this model is purely open-source and has no reliance on black-box systems, meaning that the pace of improvement in this area is likely to increase rapidly. This stands in contrast to many other approaches that do not include this key attribute.
"VideoPro achieves 50.4% accuracy on LVBench, surpassing GPT-4o and matching the performance of Qwen2.5VL-72B on VideoMME."
— VideoPro Benchmarking Results