4 papers
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
Xiaoyan Su, Peijie Dong, Zhenheng Tang +8
Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for profess…
SpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generation
Song Tang, Kaiyong Zhao, Yuliang Li +5
Automatically generating interactive 3D indoor scenes from natural language is crucial for virtual reality, gaming, and embodied AI. However, existing LLM-based approaches often su…
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
Penglei Sun, Yaoxian Song, Xiangru Zhu +7
Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…
LPZero: Language Model Zero-cost Proxy Search from Zero
Peijie Dong, Lujun Li, Xiang Liu +4
In spite of the outstanding performance, Neural Architecture Search (NAS) is criticized for massive computation. Recently, Zero-shot NAS has emerged as a promising approach by expl…