collaborators

8 papers

cs.CV2026

Position: Reasoning After Perception Means Reasoning Without Vision

Hongcheng Gao, Zihao Huang, Jingyi Tang +12

A common belief in multimodal research is that the perceptual weaknesses of vision--language models can be compensated by stronger language reasoning (e.g., chain-of-thought, in-co…

cs.LG2026

Predicting Mergeability of Parameter-Efficient Fine-Tuning Updates

Lin Tang, Wei Zhang, Jing Li +3

Low-rank adaptation (LoRA) makes it cheap to train many domain- and task-specific language model adapters, but whether two adapters can be merged is usually discovered only after b…

cs.SD2026

SonicBench: Dissecting the Physical Perception Bottleneck in Large Audio Language Models

Yirong Sun, Yanjun Chen, Xin Qiu +8

Large Audio Language Models (LALMs) excel at semantic and paralinguistic tasks, yet their ability to perceive the fundamental physical attributes of audio such as pitch, loudness,…

cs.CV2026

PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance

Siddarth Nilol Kundur Satish, Devesh Jaiswal, Hongyu Chen +1

Current video generation models produce high-quality aesthetic videos but often struggle to learn representations of real-world physics dynamics, resulting in artifacts such as unn…

cs.SD2025

BRACE: A Benchmark for Robust Audio Caption Quality Evaluation

Tianyu Guo, Hongyu Chen, Hao Liang +5

Automatic audio captioning is essential for audio understanding, enabling applications such as accessibility and content indexing. However, evaluating the quality of audio captions…

cs.CV2025

UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning

Hongyu Chen, Guangrun Wang

Chain-of-Thought (CoT) prompting improves reasoning in large language models (LLMs), but its reliance on unstructured text limits interpretability and executability in embodied tas…