collaborators

5 papers

cs.CV2025

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence

Ziang Zhang, Zehan Wang, Guanghao Zhang +5

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise mod…

cs.LG2025

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Minjie Hong, Zirun Guo, Yan Xia +4

Multimodal Large Language Models (MLLMs) are powerful at integrating diverse data, but they often struggle with complex reasoning. While Reinforcement learning (RL) can boost reaso…

cs.IR2025

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Ruofan Hu, Yan Xia, Minjie Hong +5

Multimodal large language models (MLLMs) have seen substantial progress in recent years. However, their ability to represent multimodal information in the acoustic domain remains u…

cs.CV2025

Enhancing Multimodal Unified Representations for Cross Modal Generalization

Hai Huang, Yan Xia, Shengpeng Ji +7

To enhance the interpretability of multimodal unified representations, many studies have focused on discrete unified representations. These efforts typically start with contrastive…

cs.IR2025

EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration

Minjie Hong, Yan Xia, Zehan Wang +8

Large language models (LLMs) are increasingly leveraged as foundational backbones in the development of advanced recommender systems, offering enhanced capabilities through their e…