6 papers
Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction
Weijian Chen, Weibo Yao, Yuhang Zhang +7
Scaling 3D Gaussian Splatting (3DGS) to large outdoor scenes is costly in both data acquisition and computation. Adopting panoramic images with equirectangular projection (ERP) can…
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
Zhenwei Shao, Mingyang Wang, Weijun Zhang +6
Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges f…
Multimodal OCR: Parse Anything from Documents
Handong Zheng, Yumeng Li, Kaile Zhang +22
We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus…
AirSim360: A Panoramic Simulation Platform within Drone View
Xian Ge, Yuling Pan, Yuhang Zhang +12
The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data…
dots.llm1 Technical Report
Bi Huo, Bin Tu, Cheng Qin +24
Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this…
CFTrack: Enhancing Lightweight Visual Tracking through Contrastive Learning and Feature Matching
Juntao Liang, Jun Hou, Weijun Zhang +1
Achieving both efficiency and strong discriminative ability in lightweight visual tracking is a challenge, especially on mobile and edge devices with limited computational resource…