7 papers
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
Lingjing Kong, Xin Liu, Guangyi Chen +9
Post-training pipelines that combine supervised fine-tuning (SFT) with reinforcement learning (RL) have emerged as the key recipe for transforming large language models (LLMs) into…
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
Martin Q. Ma, Willis Guo, Aditya Agrawal +4
Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and effic…
Act2See: Emergent Active Visual Perception for Video Reasoning
Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4
Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning p…
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
Shuchang Pan, Siddharth Banerjee, Dhruv Hebbar +9
Human conversation is organized by an implicit chain of thoughts that manifests as timed speech acts. Capturing this causal pathway is key to building natural full-duplex interacti…
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
Ce Zhang, Zifu Wan, Zhehan Kan +7
While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not alig…
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
Zifu Wan, Ce Zhang, Silong Yong +6
Recent Large Vision-Language Models (LVLMs) have introduced a new paradigm for understanding and reasoning about image input through textual responses. Although they have achieved…