6 papers
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Siqian Tong, Xuan Li, Chaozhuo Li +5
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, re…
AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning
Siqian Tong, Xuan Li, Yiwei Wang +5
Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained f…
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
Liyang Chen, Hongkai Chen, Yujun Cai +3
Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remai…
OptiSQL: Executable SQL Generation from Optical Tokens
Sifan Li, Hongkai Chen, Yujun Cai +3
Executable SQL generation is typically studied in text-to-SQL settings, where tables are provided as fully linearized textual schemas and contents. While effective, this formulatio…
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
Liyang Chen, Hongkai Chen, Yujun Cai +3
Video-to-Audio generation has made remarkable strides in automatically synthesizing sound for video. However, existing evaluation metrics, which focus on semantic and temporal alig…
Structured Attention Matters to Multimodal LLMs in Document Understanding
Chang Liu, Hongkai Chen, Yujun Cai +4
Document understanding remains a significant challenge for multimodal large language models (MLLMs). While previous research has primarily focused on locating evidence pages throug…