activity
20222026
most citedCross-Modal Causal Intervention for Medical Report Generation

48 citations · 66 across the 18 of their papers we have counts for

collaborators
Showing 2025Show all

7 papers · 1 filter

cs.RO2025

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

Yongjie Bai, Zhouxia Wang, Yang Liu +8

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…

cs.RO2025

AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning

Weixing Chen, Dafeng Chi, Yang Liu +7

The automated generation of layouts is vital for embodied intelligence and autonomous systems, supporting applications from virtual environment construction to home robot deploymen…

cs.CV2025

DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models

Shicheng Yin, Kaixuan Yin, Yang Liu +2

The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bott…

cs.CV2025

3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians

Zeming Wei, Junyi Lin, Yang Liu +4

3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI.…

cs.CV2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

Jingzhou Luo, Yang Liu, Weixing Chen +4

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer…

cs.LG2025

Cross-modal Causal Relation Alignment for Video Question Grounding

Weixing Chen, Yang Liu, Binglin Chen +3

Video question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG me…