collaborators

14 papers

cs.SD2026

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

Tony Alex, Wish Suharitdamrong, Sara Atito +5

Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content…

cs.CV2026

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning

Wenxi Gao, Guanxi Lu, Didi Zhu +5

Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potentia…

cs.CV2026

EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation

Tatiana Gaintseva, Akshit Achara, Gregory Slabaugh +2

Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data. On common prompts such as ``a photo of a nurse…

cs.CV2026

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

Jingyu Li, Zhe Liu, Dongnan Hu +10

World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing appro…

cs.CV2026

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

Didi Zhu, Changrui Chen, Stefanos Zafeiriou +1

When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual evidence? Correct answers can…

cs.LG2026

MidSteer: Optimal Affine Framework for Steering Generative Models

Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu +4

Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However,…