1 citations · 1 across the 5 of their papers we have counts for
5 papers
UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding
Ziya Zhou, Shangda Wu, Shenyang Xu +16
Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. H…
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Bohao Zhao, Chengrui Wei, Guangfeng Jiang +17
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive percept…
X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
Baolu Li, Jingyu Qian, Rui Guo +17
Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive…
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
Shangda Wu, Ziya Zhou, Yongyi Zang +4
We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks acr…
Unsupervised Disentanglement of Linear-Encoded Facial Semantics
Yutong Zheng, Yu-Kai Huang, Ran Tao +2
We propose a method to disentangle linear-encoded facial semantics from StyleGAN without external supervision. The method derives from linear regression and sparse representation l…