activity
20242026
collaborators

9 papers

cs.AI2026

Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty

Antonis Antoniades, Deepak Nathani, Ritam Saha +6

Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-based agents need to go beyond j…

cs.CL2026

What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning

Zhaotian Weng, Haoxuan Li, Xin Eric Wang +2

Despite the impressive performance of vision-language models (VLMs) on downstream tasks, their ability to understand and reason about causal relationships in visual inputs remains…

cs.CV2026

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14

Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stal…

cs.LG2026

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

Sophia Xiao Pu, Zhaotian Weng, Chengzhi Liu +4

Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent systems report strong reasoning…

cs.AI2026

Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing

Zhaotian Weng, Antonis Antoniades, Deepak Nathani +3

Open-ended self-improving agents can autonomously modify their own structural designs to advance their capabilities and overcome the limits of pre-defined architectures, thus reduc…

cs.AI2025

Images Speak Louder than Words: Understanding and Mitigating Bias in Vision-Language Model from a Causal Mediation Perspective

Zhaotian Weng, Zijun Gao, Jerone Andrews +1

Vision-language models (VLMs) pre-trained on extensive datasets can inadvertently learn biases by correlating gender information with specific objects or scenarios. Current methods…