collaborators

6 papers

cs.CV2026

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation

Dongsheng Wang, Dawei Su, Hui Huang

Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial reasoning capabilities. Typic…

cs.CV2026

STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

Miaoge Li, Dongsheng Wang, Zening Sun +3

Training-free zero-shot composed image retrieval models are recently gaining increasing research interest due to their generalizability and flexibility in unseen multimodal retriev…

cs.LG2026

Improving Sparse Autoencoder with Dynamic Attention

Dongsheng Wang, Jinsen Zhang, Dawei Su +1

Recently, sparse autoencoders (SAEs) have emerged as a promising technique for interpreting activations in foundation models by disentangling features into a sparse set of concepts…

cs.IR2026

RETLLM: Training and Data-Free MLLMs for Multimodal Information Retrieval

Dawei Su, Dongsheng Wang

Multimodal information retrieval (MMIR) has gained attention for its flexibility in handling text, images, or mixed queries and candidates. Recent breakthroughs in multimodal large…

cs.LG2025

LLM Empowered Prototype Learning for Zero and Few-Shot Tasks on Tabular Data

Peng Wang, Dongsheng Wang, He Zhao +3

Recent breakthroughs in large language models (LLMs) have opened the door to in-depth investigation of their potential in tabular data modeling. However, effectively utilizing adva…

cs.LG2025

Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data

Bingjie Zhang, Hongkang Li, Changlong Shi +5

Multi-task learning (MTL) concurrently trains a model on diverse task datasets to exploit common features, thereby improving overall performance across the tasks. Recent studies ha…