2 papers
cs.CV2025
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
Palaash Agrawal, Haidi Azaman, Cheston Tan
Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. Ho…
cs.LG2024
Compositional Learning of Visually-Grounded Concepts Using Reinforcement
Zijun Lin, Haidi Azaman, M Ganesh Kumar +1
Children can rapidly generalize compositionally-constructed rules to unseen test sets. On the other hand, deep reinforcement learning (RL) agents need to be trained over millions o…