6 papers
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
Lexi Pang, Liheng Zhang, Hang Ye +2
Text-to-image (T2I) diffusion models have shown remarkable success in generating high-quality images from text prompts. Recent efforts extend these models to incorporate conditiona…
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
Jianzhe Gao, Churan Wang, Weiyi Zhang +5
Medical video diagnosis involves inferring clinical decisions from dynamic tissue responses throughout examination processes. Existing methods rely on an end-to-end learning paradi…
LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation
Jiabei Xiao, Yizhou Wang, Chen Tang +3
AI Scientists have shown promising progress across multiple stages of the research pipeline, among which automatic scientific paper writing remains a formidable challenge. The Intr…
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework
Tianshu Wu, Xiangqi Kong, Yue Chen +5
Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods either rely on laborious task-spe…
Visually-grounded Humanoid Agents
Hang Ye, Xiaoxuan Ma, Fan Lu +3
Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privil…
GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
Wentao Wang, Hang Ye, Fangzhou Hong +5
Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying bod…