activity
20222025
most citedVisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis

2 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models

Shicheng Yin, Kaixuan Yin, Yang Liu +2

The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bott…

cs.CV2025

3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians

Zeming Wei, Junyi Lin, Yang Liu +4

3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI.…

cs.CV2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

Jingzhou Luo, Yang Liu, Weixing Chen +4

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer…

cs.CV2025

Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering

Kaixuan Jiang, Yang Liu, Weixing Chen +5

Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, an…

cs.CV20242 cited

VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis

Shicheng Yin, Kaixuan Yin, Weixing Chen +2

Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) are two dominant models for image analysis. While CNNs excel at extracting multi-scale features and ViTs effecti…

cs.CV2024

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

Xinshuai Song, Weixing Chen, Yang Liu +3

Existing Vision-Language Navigation (VLN) methods primarily focus on single-stage navigation, limiting their effectiveness in multi-stage and long-horizon tasks within complex and…