8 papers
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
Yuxin Lu, Jiayang Sun, Guibo Zhu +1
Diffusion-based talking head generation has achieved remarkable visual quality, yet scaling it to long-term videos remains challenging. The widely adopted chunk-wise paradigm intro…
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
Purui Bai, Tao Wu, Jiayang Sun +3
The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual…
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
Jiayang Sun, Zixin Guo, Min Cao +2
Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring…
The Vertical Challenge of Low-Altitude Economy: Why We Need a Unified Height System?
Shuaichen Yan, Xiao Hu, Jiayang Sun +5
The explosive growth of the low-altitude economy, driven by eVTOLs and UAVs, demands a unified digital infrastructure to ensure safety and scalability. However, the current aviatio…
Semantically Aware UAV Landing Site Assessment from Remote Sensing Imagery via Multimodal Large Language Models
Chunliang Hua, Zeyuan Yang, Lei Zhang +4
Safe UAV emergency landing requires more than just identifying flat terrain; it demands understanding complex semantic risks (e.g., crowds, temporary structures) invisible to tradi…
AI based signage classification for linguistic landscape studies
Yuqin Jiang, Song Jiang, Jacob Algrim +8
Linguistic Landscape (LL) research traditionally relies on manual photography and annotation of public signages to examine distribution of languages in urban space. While such meth…