4 papers
Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning
Peng, Lee, Yin Zhang +9
Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal…
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
Zishan Liu, Ruoxi Zang, Yanglin Zhang +5
Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, ofte…
Imagine a City: CityGenAgent for Procedural 3D City Generation
Zishan Liu, Zecong Tang, RuoCheng Wu +6
The automated generation of interactive 3D cities is a critical challenge with broad applications in autonomous driving, virtual reality, and embodied intelligence. While recent ad…
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
Xinzhe Zheng, Zhen-Qun Yang, Zishan Liu +4
Large Language Models (LLMs) deliver strong performance but are difficult to deploy under tight memory and compute constraints. Low-bit post-training quantization (PTQ) is a promis…