8 papers
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
Zizhuo Lin, Quanling Liu, Jinsheng Quan +6
Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gradually across turns. When a cl…
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
Kecheng Zhang, Zongxin Yang, Mingfei Han +6
Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventi…
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
Haihong Hao, Lei Chen, Mingfei Han +5
Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by action…
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
Long Li, Shuichen Ji, Ziyang Luo +4
Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify ke…
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
Longtao Jiang, Jie Huang, Mingfei Han +5
Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant,…
Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
Mingfei Han, Haihong Hao, Jinxing Zhou +5
Vision-language models often hallucinate details, generating non-existent objects or inaccurate attributes that compromise output reliability. Existing methods typically address th…