3 papers
cs.CV2026
DeepLatent: Think with Images via Parallel Latent Visual Reasoning
Dongchen Lu, Zhimo Li, Mao Shu +1
The emerging paradigm of "thinking with images" embeds visual states into intermediate reasoning steps, defining a new frontier for Vision-Language Models. Existing approaches dive…
cs.CV2025
InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression
Dongchen Lu, Yuyao Sun, Zilu Zhang +4
Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language model (LLM). However, a great qua…
cs.CV2021
OSKDet: Towards Orientation-sensitive Keypoint Localization for Rotated Object Detection
Dongchen Lu
Rotated object detection is a challenging issue of computer vision field. Loss of spatial information and confusion of parametric order have been the bottleneck for rotated detecti…