6 papers
Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes
Yifan Jiang, Cong Zhang, Bofei Zhang +4
Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and external knowledge, a synergy ove…
An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
Ni Zhang, Zhiguang Cao, Jianan Zhou +2
Complex vehicle routing problems (VRPs) remain a fundamental challenge, demanding substantial expert effort for intent interpretation and algorithm design. While large language mod…
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
Congzhi Zhang, Zhibin Wang, Yinchao Ma +5
While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasonin…
Adaptive Overclocking: Dynamic Control of Thinking Path Length via Real-Time Reasoning Signals
Shuhao Jiang, Songbo Wang, Yang Qiao +7
Large Reasoning Models (LRMs) often suffer from computational inefficiency due to overthinking, where a fixed reasoning budget fails to match the varying complexity of tasks. To ad…
The Interpretability Analysis of the Model Can Bring Improvements to the Text-to-SQL Task
Cong Zhang
To elevate the foundational capabilities and generalization prowess of the text-to-SQL model in real-world applications, we integrate model interpretability analysis with execution…
Lifelong Learner: Discovering Versatile Neural Solvers for Vehicle Routing Problems
Shaodi Feng, Zhuoyi Lin, Jianan Zhou +5
Deep learning has been extensively explored to solve vehicle routing problems (VRPs), which yields a range of data-driven neural solvers with promising outcomes. However, most neur…