8 papers
Are Video Reasoning Models Ready to Go Outside?
Yangfan He, Changgyu Boo, Jaehong Yoon
In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such conditions, their understanding and reasonin…
Confidence-Aware Tool Orchestration for Robust Video Understanding
Yangfan He, Yujin Choi, Jaehong Yoon
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: under realistic perturbations such…
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
Yi Xin, Siqi Luo, Tianxiang Xu +13
Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and ef…
Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin, Jianjiang Yang, Siqi Luo +10
Pre-trained vision models (PVMs) have demonstrated remarkable adaptability across a wide range of downstream vision tasks, showcasing exceptional performance. However, as these mod…
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
Zhongwei Wan, Zhihao Dou, Che Liu +11
Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle with complex problems requiring explicit self-reflection and self-…
MountainLion: A Multi-Modal LLM-Based Agent System for Interpretable and Adaptive Financial Trading
Siyi Wu, Junqiao Wang, Zhaoyang Guan +11
Cryptocurrency trading is a challenging task requiring the integration of heterogeneous data from multiple modalities. Traditional deep learning and reinforcement learning approach…