3 papers
cs.CV2026
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
An Lanji, Dawei Liu, Jin Li +3
Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods…
cs.CV2026
DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning
An Lanji, Dawei Liu, Jin Li +3
Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhibit unfaithfulness: the stat…
eess.SY2026
Self-Evolving Learning for Embodied AI with Criticality Model
Linxuan He, Yuying Tian, Lingxiang Fan +5
Despite rapid advances in policy pretraining, embodied AI systems routinely plateau during task-specific finetuning. The root cause lies in how finetuning data are collected: the d…