works on

From the 1 of 19 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling

Shuyuan Zhang, Chenhan Jiang, Zuoou Li +1

3D generation from natural language offers significant potential to reduce expert manual modeling efforts and enhance accessibility to 3D assets. However, existing methods often yi…

cs.CV2025

AdsQA: Towards Advertisement Video Understanding

Xinwei Long, Kai Tian, Peng Xu +10

Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpo…

cs.CV2025

Context-Aware Autoregressive Models for Multi-Conditional Image Generation

Yixiao Chen, Zhiyuan Ma, Guoli Jia +3

Autoregressive transformers have recently shown impressive image generation quality and efficiency on par with state-of-the-art diffusion models. Unlike diffusion architectures, au…

cs.CV2025

WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation

Daoan Zhang, Che Jiang, Ruoshi Xu +7

Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models still struggle with prompts that require rich world knowledge and implicit r…

cs.CV2025

JoyType: A Robust Design for Multilingual Visual Text Creation

Chao Li, Chen Jiang, Xiaolong Liu +2

Generating images with accurately represented text, especially in non-Latin languages, poses a significant challenge for diffusion models. Existing approaches, such as the integrat…

cs.CV2024

Octopus: Embodied Vision-Language Programmer from Environmental Feedback

Jingkang Yang, Yuhao Dong, Shuai Liu +8

Large vision-language models (VLMs) have achieved substantial progress in multimodal perception and reasoning. When integrated into an embodied agent, existing embodied VLM works e…