3 papers
cs.CV2026
Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation
Dongsheng Wang, Dawei Su, Hui Huang
Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial reasoning capabilities. Typic…
cs.LG2026
Improving Sparse Autoencoder with Dynamic Attention
Dongsheng Wang, Jinsen Zhang, Dawei Su +1
Recently, sparse autoencoders (SAEs) have emerged as a promising technique for interpreting activations in foundation models by disentangling features into a sparse set of concepts…
cs.IR2026
RETLLM: Training and Data-Free MLLMs for Multimodal Information Retrieval
Dawei Su, Dongsheng Wang
Multimodal information retrieval (MMIR) has gained attention for its flexibility in handling text, images, or mixed queries and candidates. Recent breakthroughs in multimodal large…