1 citations · 1 across the 3 of their papers we have counts for
4 papers
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hengyi Xie, Chenfei Yao, Xianjin Wu +4
Vision-language-action (VLA) models commonly adopt an LLM-centric pathway, where visual observations are projected into the representation space of a large language…
Generative Compositor for Few-Shot Visual Information Extraction
Zhibo Yang, Wei Hua, Sibo Song +4
Visual Information Extraction (VIE), aiming at extracting structured information from visually rich document images, plays a pivotal role in document processing. Considering variou…
Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
Linger Deng, Linghao Zhu, Yuliang Liu +6
Large Multimodal Models (LMMs) face limitations in geometric reasoning due to insufficient Chain of Thought (CoT) image-text training data. While existing approaches leverage templ…
Enhancing Scene Text Detectors with Realistic Text Image Synthesis Using Diffusion Models
Ling Fu, Zijie Wu, Yingying Zhu +2
Scene text detection techniques have garnered significant attention due to their wide-ranging applications. However, existing methods have a high demand for training data, and obta…