activity
20182026
most citedEnhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models

23 citations · 43 across the 36 of their papers we have counts for

collaborators

39 papers

cs.SD2026

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +6

Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR aro…

cs.SD2026

Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +8

An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model wri…

cs.CV2026

MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

Shuai Wang, Wangyuan Ding, Yixian Shen +5

Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with for…

cs.CV2026

Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs

Pengyu Zhang, Klim Zaporojets, Congfeng Cao +2

Multi-Modal Knowledge Graphs (MMKGs) enrich entities with multiple modalities such as text and images, yet entities with highly similar multi-modal features remain difficult to dis…

cs.LG2026

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

Yixian Shen, Zhiheng Yang, Qi Bi +6

Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substan…

cs.AI2026

A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding

Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6

Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…