3 citations · 5 across the 11 of their papers we have counts for
7 papers · 1 filter
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Zijun Lin, Zeqing Wang, Cheston Tan +2
Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are govern…
Stencil: Subject-Driven Generation with Context Guidance
Gordon Chen, Ziqi Huang, Cheston Tan +1
Recent text-to-image diffusion models can generate striking visuals from text prompts, but they often fail to maintain subject consistency across generations and contexts. One majo…
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
Zijun Lin, Shuting He, Cheston Tan +1
Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual gr…
Human-like compositional learning of visually-grounded concepts using synthetic environments
Zijun Lin, M Ganesh Kumar, Cheston Tan
The compositional structure of language enables humans to decompose complex phrases and map them to novel visual concepts, showcasing flexible intelligence. While several algorithm…
Evaluating the Generation of Spatial Relations in Text and Image Generative Models
Shang Hong Sim, Clarence Lee, Alvin Tan +1
Understanding spatial relations is a crucial cognitive ability for both humans and AI. While current research has predominantly focused on the benchmarking of text-to-image (T2I) m…
DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using Determiners
Clarence Lee, M Ganesh Kumar, Cheston Tan
State-of-the-art visual grounding models can achieve high detection accuracy, but they are not designed to distinguish between all objects versus only certain objects of interest.…