3 papers
cs.CV2026
LingDT-VL-OCR: Structure-Aware Document-Level Parsing with Fine-Grained Visual Reference
Siyi Qian, Xiongfei Bai, Bingtao Fu +4
In this paper, we propose LingDT-VL-OCR, a document parsing system tailored to financial-domain documents, transforming ultra-long financial PDFs into semantically consistent, high…
cs.RO2026
SixthSense: Task-Agnostic Proprioception-Only Whole-Body Wrench Estimation for Humanoids
Xingzhou Chen, Xiayan Xu, Yan Ning +8
Humanoid robots are entering our physical world at scale, yet as oversized toys--good at singing and dancing, but short on force-interaction capabilities for practical tasks. Bridg…
cs.CV2025
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
Yuying Li, Siyi Qian, Hao Liang +4
Geometric reasoning remains a core challenge for Multimodal Large Language Models (MLLMs). Even the most advanced closed-source systems, such as GPT-O3 and Gemini-2.5-Pro, still st…