4 papers
DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
Shuo Wang, Kai Zhang, Wenyuan Huang +5
Advancing multimodal retrieval-augmented generation (RAG) for complex document understanding presents a formidable dual dilemma of accuracy and efficiency, particularly in graph RA…
KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems
Shuo Wang, Fang Xi, Wenyuan Huang +2
Modern LLM systems increasingly rely on knowledge-selection processes that produce high-value structured priors, such as ranked evidence, graph topology, multimodal alignment, and…
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
Ting Huang, Zhenyu Zhang, Wenyuan Huang +2
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under chan…
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
Wenyuan Huang, Zhenyu Zhang, Zhao Wang +4
3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by generalization and recent zero-shot metho…