2 citations · 2 across the 5 of their papers we have counts for
10 papers
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
Baihui Xiao, Chengjian Feng, Zhijian Huang +3
Collecting real-world data for rare high-risk scenarios, long-tailed driving events, and complex interactions remains challenging, leading to poor performance of existing autonomou…
X-SAM: From Segment Anything to Any Segmentation
Hao Wang, Limeng Qiao, Zequn Jie +6
Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although…
DisTime: Distribution-based Time Representation for Video Large Language Models
Yingsen Zeng, Zepeng Huang, Yujie Zhong +4
Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and…
AP-CAP: Advancing High-Quality Data Synthesis for Animal Pose Estimation via a Controllable Image Generation Pipeline
Lei Wang, Yujie Zhong, Xiaopeng Sun +5
The task of 2D animal pose estimation plays a crucial role in advancing deep learning applications in animal behavior analysis and ecological research. Despite notable progress in…
InstructVEdit: A Holistic Approach for Instructional Video Editing
Chi Zhang, Chengjian Feng, Feng Yan +5
Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only li…
Boosting Robotic Manipulation Generalization with Minimal Costly Data
Liming Zheng, Feng Yan, Fanfan Liu +3
The growing adoption of Vision-Language-Action (VLA) models in embodied AI intensifies the demand for diverse manipulation demonstrations. However, high costs associated with data…