2 papers
cs.CV2026
AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning
Junqi Wu, Kaihua Tang, Xuanwen Chen +3
Open-world 3D affordance grounding requires localizing functional object parts in 3D given free-form language queries. Existing methods typically assume pre-built object-centric 3D…
cs.LG2026
FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning
Zekai Chen, Kairui Yang, Xuaner Chen +4
Multimodal graph foundation models aim to learn reusable knowledge from graphs enriched with text, images, attributes, and relational topology, thereby supporting diverse graph-cen…