From the 1 of 19 linked papers with an AI index.
6 papers · 1 filter
ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling
Shuyuan Zhang, Chenhan Jiang, Zuoou Li +1
3D generation from natural language offers significant potential to reduce expert manual modeling efforts and enhance accessibility to 3D assets. However, existing methods often yi…
AdsQA: Towards Advertisement Video Understanding
Xinwei Long, Kai Tian, Peng Xu +10
Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpo…
Context-Aware Autoregressive Models for Multi-Conditional Image Generation
Yixiao Chen, Zhiyuan Ma, Guoli Jia +3
Autoregressive transformers have recently shown impressive image generation quality and efficiency on par with state-of-the-art diffusion models. Unlike diffusion architectures, au…
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
Daoan Zhang, Che Jiang, Ruoshi Xu +7
Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models still struggle with prompts that require rich world knowledge and implicit r…
JoyType: A Robust Design for Multilingual Visual Text Creation
Chao Li, Chen Jiang, Xiaolong Liu +2
Generating images with accurately represented text, especially in non-Latin languages, poses a significant challenge for diffusion models. Existing approaches, such as the integrat…
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
Jingkang Yang, Yuhao Dong, Shuai Liu +8
Large vision-language models (VLMs) have achieved substantial progress in multimodal perception and reasoning. When integrated into an embodied agent, existing embodied VLM works e…