3 papers
cs.LG2026
Bottleneck Tokens for Unified Multimodal Retrieval
Siyu Sun, Jing Ren, Zhaohe Liao +8
Adapting decoder-only multimodal large language models (MLLMs) for unified multimodal retrieval faces two structural gaps. First, existing methods rely on implicit pooling, which o…
cs.CV2025
High-Quality 3D Head Reconstruction from Any Single Portrait Image
Jianfu Zhang, Yujie Gao, Jiahui Zhan +4
In this work, we introduce a novel high-fidelity 3D head reconstruction method from a single portrait image, regardless of perspective, expression, or accessories. Despite signific…
cs.CV2019
Hard Pixel Mining for Depth Privileged Semantic Segmentation
Zhangxuan Gu, Li Niu, Haohua Zhao +1
Semantic segmentation has achieved remarkable progress but remains challenging due to the complex scene, object occlusion, and so on. Some research works have attempted to use extr…