4 papers
Agentic Flow Steering and Parallel Rollout Search for Spatially Grounded Text-to-Image Generation
Ping Chen, Daoxuan Zhang, Xiangming Wang +3
Precise Text-to-Image (T2I) generation has achieved great success but is hindered by the limited relational reasoning of static text encoders and the error accumulation in open-loo…
MARO: Learning Stronger Reasoning from Social Interaction
Yin Cai, Zhouhong Gu, Juntao Zhang +1
Humans face countless scenarios that require reasoning and judgment in daily life. However, existing large language model training methods primarily allow models to learn from exis…
MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments
Yin Cai, Zhouhong Gu, Zhaohan Du +5
Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly i…
From Language To Vision: A Case Study of Text Animation
Ping Chen, Richard Alo, Justin Rundell
Information can be expressed in multiple formats including natural language, images, and motions. Human intelligence usually faces little difficulty to convert from one format to a…