23 citations · 43 across the 36 of their papers we have counts for
39 papers
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +6
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR aro…
Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar +8
An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model wri…
MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding
Shuai Wang, Wangyuan Ding, Yixian Shen +5
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with for…
Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs
Pengyu Zhang, Klim Zaporojets, Congfeng Cao +2
Multi-Modal Knowledge Graphs (MMKGs) enrich entities with multiple modalities such as text and images, yet entities with highly similar multi-modal features remain difficult to dis…
Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning
Yixian Shen, Zhiheng Yang, Qi Bi +6
Multimodal spatial reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substan…
A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
Shuai Wang, Hongyi Zhu, Jia-Hong Huang +6
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise…