collaborators

12 papers

cs.CV2026

MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

Shuai Wang, Wangyuan Ding, Yixian Shen +5

Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with for…

cs.CV2026

Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs

Pengyu Zhang, Klim Zaporojets, Congfeng Cao +2

Multi-Modal Knowledge Graphs (MMKGs) enrich entities with multiple modalities such as text and images, yet entities with highly similar multi-modal features remain difficult to dis…

cs.CV2026

Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding

Jiaqi Li, Shuntian Zheng, Yixian Shen +4

Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long, untrimmed videos, making video-language-model prohibitively expensive. While rec…

cs.LG2026

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

Yixian Shen, Zhiheng Yang, Qi Bi +6

Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substan…

cs.AI2026

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6

Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…

cs.SD2026

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech

Jia-Hong Huang, Seulgi Kim, Yi Chieh Liu +5

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker id…