papers

Publications (59)

cs.LG2025

Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs

Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer +1

The rise of reinforcement learning (RL) in critical real-world applications demands a fundamental rethinking of privacy in AI systems. Traditional privacy frameworks, designed to p…

cs.LG2024

Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios

Shantanu Jaiswal, Debaditya Roy, Basura Fernando +1

Complex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the imm…

cs.NE2024

One-shot learning of paired association navigation with biologically plausible schemas

M Ganesh Kumar, Cheston Tan, Camilo Libedinsky +2

Schemas are knowledge structures that can enable rapid learning. Rodent one-shot learning in a multiple paired association navigation task has been postulated to be schema-dependen…

cs.LG2025

FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF

Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong +2

In the era of increasing privacy concerns and demand for personalized experiences, traditional Reinforcement Learning with Human Feedback (RLHF) frameworks face significant challen…

cs.CV2025

Stencil: Subject-Driven Generation with Context Guidance

Gordon Chen, Ziqi Huang, Cheston Tan +1

Recent text-to-image diffusion models can generate striking visuals from text prompts, but they often fail to maintain subject consistency across generations and contexts. One majo…

cs.LG2024

CAESAR: Enhancing Federated RL in Heterogeneous MDPs through Convergence-Aware Sampling with Screening

Hei Yi Mak, Flint Xiaofeng Fan, Luca A. Lanzendörfer +3

In this study, we delve into Federated Reinforcement Learning (FedRL) in the context of value-based agents operating across diverse Markov Decision Processes (MDPs). Existing FedRL…

cs.CL2024

LLM-Based Multi-Hop Question Answering with Knowledge Graph Integration in Evolving Environments

Ruirui Chen, Weifeng Jiang, Chengwei Qin +5

The important challenge of keeping knowledge in Large Language Models (LLMs) up-to-date has led to the development of various methods for incorporating new facts. However, existing…

cs.LG2015

Deep Convolutional Networks are Hierarchical Kernel Machines

Fabio Anselmi, Lorenzo Rosasco, Cheston Tan +1

In i-theory a typical layer of a hierarchical architecture consists of HW modules pooling the dot products of the inputs to the layer with the transformations of a few templates un…

cs.RO2026

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

Zijun Lin, Jiafei Duan, Haoquan Fang +4

Recent advances in robotic manipulation have integrated low-level robotic control into Vision-Language Models (VLMs), extending them into Vision-Language-Action (VLA) models. Altho…

cs.LG2022

A Survey on Machine Learning Approaches for Modelling Intuitive Physics

Jiafei Duan, Arijit Dasgupta, Jason Fischer +1

Research in cognitive science has provided extensive evidence of human cognitive ability in performing physical reasoning of objects from noisy perceptual inputs. Such a cognitive…

cs.AI2026

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs

Kevin Wang, Anna Thöni, Benjamin Kempinski +50

Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly underst…

cs.AI2025

The Singapore Consensus on Global AI Safety Research Priorities

Yoshua Bengio, Tegan Maharaj, Luke Ong +84

Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…

cs.CL2026

CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?

Ruirui Chen, Weifeng Jiang, Chengwei Qin +1

Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Large Language Models (LLMs) become ubiqu…

cs.CV2023

Read My Mind: A Multi-Modal Dataset for Human Belief Prediction

Jiafei Duan, Samson Yu, Nicholas Tan +2

Understanding human intentions is key to enabling effective and efficient human-robot interaction (HRI) in collaborative settings. To enable developments and evaluation of the abil…

cs.AI2021

SPACE: A Simulator for Physical Interactions and Causal Learning in 3D Environments

Jiafei Duan, Samson Yu Bai Jian, Cheston Tan

Recent advancements in deep learning, computer vision, and embodied AI have given rise to synthetic causal reasoning video datasets. These datasets facilitate the development of AI…

q-bio.NC2017

Neurogenesis and multiple plasticity mechanisms enhance associative memory retrieval in a spiking network model of the hippocampus

Yansong Chua, Cheston Tan

Hippocampal CA3 is crucial for the formation of long-term associative memory. It has a heavily recurrent connectivity, and memories are thought to be stored as memory engrams in th…

cs.CV2024

Evaluating the Generation of Spatial Relations in Text and Image Generative Models

Shang Hong Sim, Clarence Lee, Alvin Tan +1

Understanding spatial relations is a crucial cognitive ability for both humans and AI. While current research has predominantly focused on the benchmarking of text-to-image (T2I) m…

cs.CV2021

A Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event Categories

Arijit Dasgupta, Jiafei Duan, Marcelo H. Ang +4

Recent work in computer vision and cognitive reasoning has given rise to an increasing adoption of the Violation-of-Expectation (VoE) paradigm in synthetic datasets. Inspired by in…

cs.NE2021

A nonlinear hidden layer enables actor-critic agents to learn multiple paired association navigation

M Ganesh Kumar, Cheston Tan, Camilo Libedinsky +2

Navigation to multiple cued reward locations has been increasingly used to study rodent learning. Though deep reinforcement learning agents have been shown to be able to learn the…

cs.CV2021

AVoE: A Synthetic 3D Dataset on Understanding Violation of Expectation for Artificial Cognition

Arijit Dasgupta, Jiafei Duan, Marcelo H. Ang +1

Recent work in cognitive reasoning and computer vision has engendered an increasing popularity for the Violation-of-Expectation (VoE) paradigm in synthetic datasets. Inspired by wo…

cs.LG2023

Good Time to Ask: A Learning Framework for Asking for Help in Embodied Visual Navigation

Jenny Zhang, Samson Yu, Jiafei Duan +1

In reality, it is often more efficient to ask for help than to search the entire space to find an object with an unknown location. We present a learning framework that enables an a…

cs.AI2014

Neural tuning size is a key factor underlying holistic face processing

Cheston Tan, Tomaso Poggio

Faces are a class of visual stimuli with unique significance, for a variety of reasons. They are ubiquitous throughout the course of a person's life, and face recognition is crucia…

cs.CV2025

Human-like compositional learning of visually-grounded concepts using synthetic environments

Zijun Lin, M Ganesh Kumar, Cheston Tan

The compositional structure of language enables humans to decompose complex phrases and map them to novel visual concepts, showcasing flexible intelligence. While several algorithm…

cs.AI2023

Advancing Perception in Artificial Intelligence through Principles of Cognitive Science

Palaash Agrawal, Cheston Tan, Heena Rathore

Although artificial intelligence (AI) has achieved many feats at a rapid pace, there still exist open problems and fundamental shortcomings related to performance and resource effi…

cs.RO2025

10 Open Challenges Steering the Future of Vision-Language-Action Models

Soujanya Poria, Navonil Majumder, Chia-Yu Hung +7

Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly prevalent in the embodied AI arena, following the widespread succ…

cs.LG2024

Social Learning through Interactions with Other Agents: A Survey

Dylan Hillier, Cheston Tan, Jing Jiang

Social learning plays an important role in the development of human intelligence. As children, we imitate our parents' speech patterns until we are able to produce sounds; we learn…

cs.AI2026

TangramSR: Can Vision-Language Models Reason in Continuous Geometric Space?

Yikun Zong, Cheston Tan

Humans excel at spatial reasoning tasks like Tangram puzzle assembly through cognitive processes involving mental rotation, iterative refinement, and visual feedback. Inspired by h…

cs.CV2024

Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal +2

While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the…

cs.AI2022

ABCDE: An Agent-Based Cognitive Development Environment

Jieyi Ye, Jiafei Duan, Samson Yu +2

Children's cognitive abilities are sometimes cited as AI benchmarks. How can the most common 1,000 concepts (89\% of everyday use) be learnt in a naturalistic children's setting? C…

cs.AI2026

MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games

Yunfei Xie, Kevin Wang, Bobby Cheng +9

Multi-turn, multi-agent LLM game evaluations often exhibit substantial run-to-run variance. In long-horizon interactions, small early deviations compound across turns and are ampli…

cs.CV2025

STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning

Palaash Agrawal, Haidi Azaman, Cheston Tan

Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. Ho…

cs.LG2022

Fault-Tolerant Federated Reinforcement Learning with Theoretical Guarantee

Flint Xiaofeng Fan, Yining Ma, Zhongxiang Dai +3

The growing literature of Federated Learning (FL) has recently inspired Federated Reinforcement Learning (FRL) to encourage multiple agents to federatively build a better decision-…

cs.CV2022

TDAM: Top-Down Attention Module for Contextually Guided Feature Selection in CNNs

Shantanu Jaiswal, Basura Fernando, Cheston Tan

Attention modules for Convolutional Neural Networks (CNNs) are an effective method to enhance performance on multiple computer-vision tasks. While existing methods appropriately mo…

cs.CL2024

Can LLMs perform structured graph reasoning?

Palaash Agrawal, Shavak Vasania, Cheston Tan

Pretrained Large Language Models (LLMs) have demonstrated various reasoning capabilities through language-based prompts alone, particularly in unstructured task settings (tasks pur…

cs.RO2019

Efficient Robotic Task Generalization Using Deep Model Fusion Reinforcement Learning

Tianying Wang, Hao Zhang, Wei Qi Toh +5

Learning-based methods have been used to pro-gram robotic tasks in recent years. However, extensive training is usually required not only for the initial task learning but also for…

cs.CV2025

Inferring Past Human Actions in Homes with Abductive Reasoning

Clement Tan, Chai Kiat Yeo, Cheston Tan +1

Abductive reasoning aims to make the most likely inference for a given set of incomplete observations. In this paper, we introduce "Abductive Past Action Inference", a novel resear…

cs.CV2022

BOSS: A Benchmark for Human Belief Prediction in Object-context Scenarios

Jiafei Duan, Samson Yu, Nicholas Tan +2

Humans with an average level of social cognition can infer the beliefs of others based solely on the nonverbal communication signals (e.g. gaze, gesture, pose and contextual inform…

cs.CV2019

An End-to-End Network for Generating Social Relationship Graphs

Arushi Goel, Keng Teck Ma, Cheston Tan

Socially-intelligent agents are of growing interest in artificial intelligence. To this end, we need systems that can understand social relationships in diverse social contexts. In…

cs.AI2026

SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning

Bo Liu, Leon Guertler, Simon Yu +9

Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approache…

cs.LG2023

FedHQL: Federated Heterogeneous Q-Learning

Flint Xiaofeng Fan, Yining Ma, Zhongxiang Dai +3

Federated Reinforcement Learning (FedRL) encourages distributed agents to learn collectively from each other's experience to improve their performance without exchanging their raw…

cs.LG2025

FedHPD: Heterogeneous Federated Reinforcement Learning via Policy Distillation

Wenzheng Jiang, Ji Wang, Xiongtao Zhang +3

Federated Reinforcement Learning (FedRL) improves sample efficiency while preserving privacy; however, most existing studies assume homogeneous agents, limiting its applicability i…

cs.LG2024

Compositional Learning of Visually-Grounded Concepts Using Reinforcement

Zijun Lin, Haidi Azaman, M Ganesh Kumar +1

Children can rapidly generalize compositionally-constructed rules to unseen test sets. On the other hand, deep reinforcement learning (RL) agents need to be trained over millions o…

cs.CV2023

DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using Determiners

Clarence Lee, M Ganesh Kumar, Cheston Tan

State-of-the-art visual grounding models can achieve high detection accuracy, but they are not designed to distinguish between all objects versus only certain objects of interest.…

cs.RO2024

RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing

Bo Ai, Stephen Tian, Haochen Shi +4

Tactile feedback is critical for understanding the dynamics of both rigid and deformable objects in many manipulation tasks, such as non-prehensile manipulation and dense packing.…

cs.CV2021

6D Pose Estimation with Correlation Fusion

Yi Cheng, Hongyuan Zhu, Ying Sun +6

6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illum…

cs.AI2022

A Survey of Embodied AI: From Simulators to Research Tasks

Jiafei Duan, Samson Yu, Hui Li Tan +2

There has been an emerging paradigm shift from the era of "internet AI" to "embodied AI", where AI algorithms and agents no longer learn from datasets of images, videos or text cur…

cs.CL2025

Theory of Mind in Large Language Models: Assessment and Enhancement

Ruirui Chen, Weifeng Jiang, Chengwei Qin +1

Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Large Language Models (LLMs) become incre…

cs.CV2021

PIP: Physical Interaction Prediction via Mental Simulation with Span Selection

Jiafei Duan, Samson Yu, Soujanya Poria +2

Accurate prediction of physical interaction outcomes is a crucial component of human intelligence and is important for safe and efficient deployments of robots in the real world. W…

cs.CL2025

TextArena

Leon Guertler, Bobby Cheng, Simon Yu +3

TextArena is an open-source collection of competitive text-based games for training and evaluation of agentic behavior in Large Language Models (LLMs). It spans 57+ unique environm…

cs.CV2025

GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding

Zijun Lin, Shuting He, Cheston Tan +1

Sequential grounding in 3D point clouds (SG3D) refers to locating sequences of objects by following text instructions for a daily activity with detailed steps. Current 3D visual gr…

cs.CL2025

How do Transformer Embeddings Represent Compositions? A Functional Analysis

Aishik Nagar, Ishaan Singh Rawal, Mansi Dhanania +1

Compositionality is a key aspect of human intelligence, essential for reasoning and generalization. While transformer-based models have become the de facto standard for many langua…

cs.AI2026

Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol

Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer +1

As AI agents powered by large language models (LLMs) increasingly use external tools for high-stakes decisions, a critical reliability question arises: how do errors propagate acro…

cs.CL2024

Super Tiny Language Models

Dylan Hillier, Leon Guertler, Cheston Tan +3

The rapid advancement of large language models (LLMs) has led to significant improvements in natural language processing but also poses challenges due to their high computational a…

cs.CV2020

Actionet: An Interactive End-To-End Platform For Task-Based Data Collection And Augmentation In 3D Environment

Jiafei Duan, Samson Yu, Hui Li Tan +1

The problem of task planning for artificial agents remains largely unsolved. While there has been increasing interest in data-driven approaches for the study of task planning for a…

cs.CL2024

Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis

Aishik Nagar, Shantanu Jaiswal, Cheston Tan

Vision-language models (VLMs) have shown impressive zero- and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visua…

cs.LG2023

Robustness of Utilizing Feedback in Embodied Visual Navigation

Jenny Zhang, Samson Yu, Jiafei Duan +1

This paper presents a framework for training an agent to actively request help in object-goal navigation tasks, with feedback indicating the location of the target object in its fi…

cs.CV2026

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

Zijun Lin, Zeqing Wang, Cheston Tan +2

The paper introduces StatePlay, a game world model that jointly predicts visual frames and internal game states using a mixture-of-transformers architecture to generate gameplay th…

#game world modeling#state-aware generation#mechanics consistency#cross-modal interaction
cs.AI2025

From Grunts to Lexicons: Emergent Language from Cooperative Foraging

Maytus Piriyajitakonkij, Rujikorn Charakorn, Weicheng Tao +4

Language is a powerful communicative and cognitive tool. It enables humans to express thoughts, share intentions, and reason about complex phenomena. Despite our fluency in using a…

cs.CL2024

STLM Engineering Report: Dropout

Dylan Hillier, Leon Guertler, Bobby Cheng +1

In this work we explore the relevance of dropout for modern language models, particularly in the context of models on the scale of <100M parameters. We explore it's relevance first…