2 papers
cs.CV2026
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
Dongchuan Ran, Linyu Ou, Xueheng Li +5
Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks ad…
cs.CV2025
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
Linyu Ou, YuYang Yin
While Reinforcement Learning with Verifiable Rewards has enhanced the reasoning of large-scale language models (LLMs), its efficacy for lightweight multimodal language models (MLLM…