4 papers
Hermite Curves as Trajectory Priors for Vision-Language-Action Models
Qi Lv, Jianming Xing, Zhao Yang +5
Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Existing work typically flatten eac…
ChunkFlow: Towards Continuity-Consistent Chunked Policy Learning
Zhao Yang, Yinan Shi, Mingyuan Yao +3
Vision-language action (VLA) models increasingly adopt chunked action heads to satisfy real-time constraints; however, this introduces boundary jitter: overlapping regions between…
CityGen: Structure-Guided City-Style Synthesis for Cross-City Autonomous Driving
Zezhong Qian, Zhao Yang, Lu Tan +4
Autonomous driving systems are commonly trained and evaluated within limited geographic regions, which hinders their scalability when deployed in new cities. However, significant d…
MetaFood CVPR 2024 Challenge on Physically Informed 3D Food Reconstruction: Methods and Results
Jiangpeng He, Yuhao Chen, Gautham Vinod +16
The increasing interest in computer vision applications for nutrition and dietary monitoring has led to the development of advanced 3D reconstruction techniques for food items. How…