7 papers
ScalingAR: Scaling Confidence for Autoregressive Image Generation
Harold Haodong Chen, Xianfeng Wu, Wen-Jie Shu +4
Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregressive (AR) image generation rem…
Panoramic Affordance Prediction
Zixin Zhang, Chenfei Liao, Hongfei Zhang +10
Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from n…
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
Harold Haodong Chen, Xinxiang Yin, Wen-Jie Shu +6
Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human…
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
Harold Haodong Chen, Disen Lan, Wen-Jie Shu +10
The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consi…
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
Zixin Zhang, Kanghao Chen, Hanqing Wang +5
Affordance prediction, which identifies interaction regions on objects based on language instructions, is critical for embodied AI. Prevailing end-to-end models couple high-level r…
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
Hongfei Zhang, Kanghao Chen, Zixin Zhang +6
This paper presents DualCamCtrl, a novel end-to-end diffusion model for camera-controlled video generation. Recent works have advanced this field by representing camera poses as ra…