From the 1 of 9 linked papers with an AI index.
9 papers
HierCAD: Hierarchical Text-to-CAD Design via Structure Alignment and Parameter Grounding
Jimin Xu, Tianbao Wang, Tao Jin +1
HierCAD is a system that turns natural language descriptions into CAD models by breaking the design process into hierarchical steps and aligning the generated geometry with its und…
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
Ruofan Hu, Menghui Zhu, Jieming Zhu +8
Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typically combine dense visual e…
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
Zequn Xie, Guijin Luo, Chuxin Wang +4
Text-based person anomaly search retrieves specific behavioral events from surveillance archives using natural-language queries. Although recent pose-aware methods align geometric…
From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents
Rongsheng Zhang, Ruofan Hu, Weijie Chen +7
While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Current systems typically rely…
ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
Jiayang Xu, Fan Zhuo, Majun Zhang +6
Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks can be formulated as a decoup…
From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning
Xiaoda Yang, Yuxiang Liu, Shenzhou Gao +8
Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A…