9 papers
Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving
Meibo Hu, Jiamian Wang, Pichao Wang +1
Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs and high-level navigation in…
Attention-Steered Vision-Language Models for Sign Language Translation
Meibo Hu, Guohao Sun, Annemarie D. Ross +2
Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we…
Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage
Vivek Senthil, Zhiqiang Tao, Ernest Fokoué
Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability o…
Information-Regularized Attention for Visual-Centric Reasoning
Guohao Sun, Xiaofang Wang, Yash Patel +3
Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting af…
DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents
Jiamian Wang, Ruiyi Zhang, Tong Yu +5
Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples serve as the training envir…
Latent Chain-of-Thought for Visual Reasoning
Guohao Sun, Hang Hua, Jian Wang +5
Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…