most citedUnsupervised Disentanglement of Linear-Encoded Facial Semantics

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2026

UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

Ziya Zhou, Shangda Wu, Shenyang Xu +16

Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. H…

cs.CV2026

X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving

Bohao Zhao, Chengrui Wei, Guangfeng Jiang +17

Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive percept…

cs.CV2026

X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

Baolu Li, Jingyu Qian, Rui Guo +17

Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive…

cs.SD2026

Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding

Shangda Wu, Ziya Zhou, Yongyi Zang +4

We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks acr…

cs.CV2021★ 1 cited

Unsupervised Disentanglement of Linear-Encoded Facial Semantics

Yutong Zheng, Yu-Kai Huang, Ran Tao +2

We propose a method to disentangle linear-encoded facial semantics from StyleGAN without external supervision. The method derives from linear regression and sparse representation l…