activity
20242026
most citedHOI-aware Adaptive Network for Weakly-supervised Action Segmentation

7 citations · 7 across the 2 of their papers we have counts for

collaborators

6 papers

cs.RO2026

CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos

Chubin Zhang, Jianan Wang, Zifeng Gao +5

Generalist Vision-Language-Action models remain constrained by the scarcity of robotic data relative to the abundance of human video demonstrations. Existing Latent Action Models a…

cs.CV20267 cited

HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

Runzhong Zhang, Suchen Wang, Yueqi Duan +3

In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of…

cs.CV2025

Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline

Linqing Zhao, Xiuwei Xu, Yirui Wang +5

Incrementally recovering real-sized 3D geometry from a pose-free RGB stream is a challenging task in 3D reconstruction, requiring minimal assumptions on input data. Existing method…

cs.CV2025

Q-VLM: Post-training Quantization for Large Vision-Language Models

Changyuan Wang, Ziwei Wang, Xiuwei Xu +3

In this paper, we propose a post-training quantization framework of large vision-language models (LVLMs) for efficient multi-modal inference. Conventional quantization methods sequ…

cs.CV2024

Learning Dual-Level Deformable Implicit Representation for Real-World Scale Arbitrary Super-Resolution

Zhiheng Li, Muheng Li, Jixuan Fan +4

Scale arbitrary super-resolution based on implicit image function gains increasing popularity since it can better represent the visual world in a continuous manner. However, existi…

cs.CV2024

GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian Generation

Chubin Zhang, Hongliang Song, Yi Wei +3

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only…