12 papers
MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding
Shuai Wang, Wangyuan Ding, Yixian Shen +5
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with for…
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
Jiaqi Li, Shuntian Zheng, Yixian Shen +4
Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long, untrimmed videos, making video-language-model prohibitively expensive. While rec…
Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning
Yixian Shen, Zhiheng Yang, Qi Bi +6
Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substan…
Fast Transformer Inference on ARM-Based HMPSoCs
Hang Xu, Yixian Shen, Thanassis Giannetsos +1
Transformer models have set new performance standards for machine learning (ML) tasks. However, their resource-intensive deployment on resource-constrained edge devices for cloud-f…
A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
Jia-Hong Huang, Seulgi Kim, Yi Chieh Liu +5
Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker id…