17 citations · 18 across the 3 of their papers we have counts for
4 papers
Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
Ze Feng, Jiang-jiang Liu, Sen Yang +5
The computational expense of redundant vision tokens in Large Vision-Language Models (LVLMs) has led many existing methods to compress them via a vision projector. However, this co…
Attend to Who You Are: Supervising Self-Attention for Keypoint Detection and Instance-Aware Association
Sen Yang, Zhicheng Wang, Ze Chen +7
This paper presents a new method to solve keypoint detection and instance association by using Transformer. For bottom-up multi-person pose estimation models, they need to detect k…
SIENet: Spatial Information Enhancement Network for 3D Object Detection from Point Cloud
Ziyu Li, Yuncong Yao, Zhibin Quan +2
LiDAR-based 3D object detection pushes forward an immense influence on autonomous vehicles. Due to the limitation of the intrinsic properties of LiDAR, fewer points are collected a…
TransPose: Keypoint Localization via Transformer
Sen Yang, Zhibin Quan, Mu Nie +1
While CNN-based models have made remarkable progress on human pose estimation, what spatial dependencies they capture to localize keypoints remains unclear. In this work, we propos…