14 papers
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
Tony Alex, Wish Suharitdamrong, Sara Atito +5
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content…
Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning
Wenxi Gao, Guanxi Lu, Didi Zhu +5
Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potentia…
EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
Tatiana Gaintseva, Akshit Achara, Gregory Slabaugh +2
Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data. On common prompts such as ``a photo of a nurse…
Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation
Jingyu Li, Zhe Liu, Dongnan Hu +10
World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing appro…
VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?
Didi Zhu, Changrui Chen, Stefanos Zafeiriou +1
When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence? Correct answers can…
MidSteer: Optimal Affine Framework for Steering Generative Models
Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu +4
Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However,…