5 papers · 1 filter
LogiShot: Logically Coherent Cross-Shot Video Generation
Shuai Guo, Yuhang Yang, Zeyu Zhang +4
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production,…
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Chengjun Yu, Xuhan Zhu, Chaoqun Du +4
Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-ter…
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu +10
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…
StreamingClaw Technical Report
Jiawei Chen, Zhe Chen, Chaoqun Du +21
Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing st…
LDGen: Enhancing Text-to-Image Synthesis via Large Language Model-Driven Language Representation
Pengzhi Li, Pengfei Yu, Zide Liu +5
In this paper, we introduce LDGen, a novel method for integrating large language models (LLMs) into existing text-to-image diffusion models while minimizing computational demands.…