4 papers
Agentic Very Long Video Understanding
Aniket Rege, Arka Sadhu, Yuliang Li +5
The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond sho…
PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models
Yuliang Li, Chu Zhou, Heng Guo +3
Mainstream vision-language models (VLMs) fundamentally struggle with severe optical ambiguities, such as reflections and transparent objects, due to the inherent limitations of sta…
SpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generation
Song Tang, Kaiyong Zhao, Yuliang Li +5
Automatically generating interactive 3D indoor scenes from natural language is crucial for virtual reality, gaming, and embodied AI. However, existing LLM-based approaches often su…
Text-to-Stage: Spatial Layouts from Long-form Narratives
Jefferson Hernandez, Swarnadeep Saha, Chenxi Whitehouse +6
In this work, we probe the ability of a language model to demonstrate spatial reasoning from unstructured text, mimicking human capabilities and automating a process that benefits…