4 papers
Video-HOCA: A Diagnostic Benchmark for Physical Anomaly Reasoning in Video-LLMs
Chang Liu, Yunfan Ye, Qingyang Zhou +5
We introduce Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos. Video-HOCA uses an Ontological-Causal taxonomy to distinguish violations of an entity's ow…
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
Chang Liu, Yunfan Ye, Fan Zhang +3
Numerous synthesized videos from generative models, especially human-centric ones that simulate realistic human actions, pose significant threats to human information security and…
Learning Cross-hand Policies for High-DOF Reaching and Grasping
Qijin She, Shishun Zhang, Yunfan Ye +2
Reaching-and-grasping is a fundamental skill for robotic manipulation, but existing methods usually train models on a specific gripper and cannot be reused on another gripper. In t…
ROICtrl: Boosting Instance Control for Visual Generation
Yuchao Gu, Yipin Zhou, Yunfan Ye +5
Natural language often struggles to accurately associate positional and attribute information with multiple instances, which limits current text-based visual generation models to s…