1 citations · 3 across the 10 of their papers we have counts for
11 papers · 1 filter
DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
Yuchuan Wu, Minghan Zhuo, Teng Fu +3
Document understanding with multimodal large language models (MLLMs) requires not only accurate answers but also explicit, evidence-grounded reasoning, especially in high-stakes sc…
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
Teng Fu, Mengyang Zhao, Ke Niu +2
LVLMs have been shown to perform excellently in image-level tasks such as VQA and caption. However, in many instance-level tasks, such as visual grounding and object detection, LVL…
Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
Kaixin Peng, Mengyang Zhao, Haiyang Yu +2
As the oldest mature writing system, Oracle Bone Script (OBS) has long posed significant challenges for archaeological decipherment due to its rarity, abstractness, and pictographi…
IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
Mengyang Zhao, Teng Fu, Haiyang Yu +2
Few-Shot Industrial Anomaly Detection (FS-IAD) has important applications in automating industrial quality inspection. Recently, some FS-IAD methods based on Large Vision-Language…
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
Teng Fu, Yuwen Chen, Zhuofan Chen +3
Multi-object tracking is a classic field in computer vision. Among them, pedestrian tracking has extremely high application value and has become the most popular research category.…
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
Ke Niu, Zhuofan Chen, Haiyang Yu +5
Computer-Aided Design (CAD) plays a pivotal role in industrial manufacturing. Orthographic projection reasoning underpins the entire CAD workflow, encompassing design, manufacturin…