5 papers
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation
Jingqi Tian, Yiheng Du, Haoji Zhang +6
Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existing methods often struggle with m…
FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis
Ruihao Xu, Xingming Shui, Jingxuan Niu +4
As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evidence and traceable accountabili…
VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
Yuji Wang, Wenlong Liu, Jingxuan Niu +2
Tool-integrated visual reasoning (TiVR) has demonstrated great potential in enhancing multimodal problem-solving. However, existing TiVR paradigms mainly focus on integrating vario…
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
Xin Gu, Haoji Zhang, Qihang Fan +7
Spatio-temporal video grounding (STVG) requires localizing a target object in untrimmed videos both temporally and spatially from natural language descriptions. Despite their stron…
InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
Yixuan Zhu, Haolin Wang, Ao Li +6
Image enhancement finds wide-ranging applications in real-world scenarios due to complex environments and the inherent limitations of imaging devices. Recent diffusion-based method…