3 citations · 3 across the 18 of their papers we have counts for
4 papers · 2 filters
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
Linpeng Huang, Weixing Chen, Zexin Chen +2
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predomin…
When Preference Labels Fall Short: Aligning Diffusion Models from Real Data
Weiyan Chen, Weijian Deng, Yao Xiao +5
Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on prefere…
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
Shicheng Yin, Kaixuan Yin, Weixing Chen +3
World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformerbased models significantly hinders real-tim…