3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2025
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Dongchen Lu, Yuyao Sun, Zilu Zhang +4
Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language model (LLM). However, a great qua…
cs.CV2024★ 3 cited
UniLoc: Towards Universal Place Recognition Using Any Single Modality
Yan Xia, Zhendong Li, Yun-Jin Li +4
To date, most place recognition methods focus on single-modality retrieval. While they perform well in specific environments, cross-modal methods offer greater flexibility by allow…