Publications (8)
Compositional Learning of Visually-Grounded Concepts Using Reinforcement
Zijun Lin, Haidi Azaman, M Ganesh Kumar +1
Children can rapidly generalize compositionally-constructed rules to unseen test sets. On the other hand, deep reinforcement learning (RL) agents need to be trained over millions o…
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
Zijun Lin, Shuting He, Cheston Tan +1
Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual gr…
All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
Lik-Hang Lee, Tristan Braud, Pengyuan Zhou +6
Since the popularisation of the Internet in the 1990s, the cyberspace has kept evolving. We have created various computer-mediated virtual environments including social networks, v…
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Zijun Lin, Zeqing Wang, Cheston Tan +2
The paper introduces StatePlay, a game world model that jointly predicts visual frames and internal game states using a mixture-of-transformers architecture to generate gameplay th…
When Creators Meet the Metaverse: A Survey on Computational Arts
Lik-Hang Lee, Zijun Lin, Rui Hu +5
The metaverse, enormous virtual-physical cyberspace, has brought unprecedented opportunities for artists to blend every corner of our physical surroundings with digital creativity.…
Human-like compositional learning of visually-grounded concepts using synthetic environments
Zijun Lin, M Ganesh Kumar, Cheston Tan
The compositional structure of language enables humans to decompose complex phrases and map them to novel visual concepts, showcasing flexible intelligence. While several algorithm…
FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction Following
Zijun Lin, Chao Tang, Hanjing Ye +1
Robotic instruction following tasks require seamless integration of visual perception, task planning, target localization, and motion execution. However, existing task planning met…
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
Zijun Lin, Jiafei Duan, Haoquan Fang +4
Recent advances in robotic manipulation have integrated low-level robotic control into Vision-Language Models (VLMs), extending them into Vision-Language-Action (VLA) models. Altho…