2 citations · 2 across the 3 of their papers we have counts for
9 papers
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
Weixing Chen, Dafeng Chi, Yang Liu +7
The automated generation of layouts is vital for embodied intelligence and autonomous systems, supporting applications from virtual environment construction to home robot deploymen…
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
Shicheng Yin, Kaixuan Yin, Yang Liu +2
The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bott…
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
Zeming Wei, Junyi Lin, Yang Liu +4
3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI.…
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
Jingzhou Luo, Yang Liu, Weixing Chen +4
3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer…
Cross-modal Causal Relation Alignment for Video Question Grounding
Weixing Chen, Yang Liu, Binglin Chen +3
Video question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG me…
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
Kaixuan Jiang, Yang Liu, Weixing Chen +5
Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, an…