activity
20232026
most citedAdvancing Perception in Artificial Intelligence through Principles of Cognitive Science

3 citations · 5 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

Zijun Lin, Zeqing Wang, Cheston Tan +2

Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are govern…

cs.CV2025

Stencil: Subject-Driven Generation with Context Guidance

Gordon Chen, Ziqi Huang, Cheston Tan +1

Recent text-to-image diffusion models can generate striking visuals from text prompts, but they often fail to maintain subject consistency across generations and contexts. One majo…

cs.CV2025

GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding

Zijun Lin, Shuting He, Cheston Tan +1

Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual gr…

cs.CV2025

Human-like compositional learning of visually-grounded concepts using synthetic environments

Zijun Lin, M Ganesh Kumar, Cheston Tan

The compositional structure of language enables humans to decompose complex phrases and map them to novel visual concepts, showcasing flexible intelligence. While several algorithm…

cs.CV20241 cited

Evaluating the Generation of Spatial Relations in Text and Image Generative Models

Shang Hong Sim, Clarence Lee, Alvin Tan +1

Understanding spatial relations is a crucial cognitive ability for both humans and AI. While current research has predominantly focused on the benchmarking of text-to-image (T2I) m…

cs.CV2023

DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using Determiners

Clarence Lee, M Ganesh Kumar, Cheston Tan

State-of-the-art visual grounding models can achieve high detection accuracy, but they are not designed to distinguish between all objects versus only certain objects of interest.…