3 papers
cs.CV2026
OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction
Taiting Lu, Runze Liu, Ziwei Dong +18
Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the…
cs.CV2026
OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing
Taiting Lu, Kaiyuan Lin, Ziwei Dong +18
Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However, their ability to reason abou…
cs.IR2026
DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
Shuo Wang, Kai Zhang, Wenyuan Huang +5
Advancing multimodal retrieval-augmented generation (RAG) for complex document understanding presents a formidable dual dilemma of accuracy and efficiency, particularly in graph RA…