5 papers
Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts
Philip Xu
This paper introduces a novel Multi-Agent Cooperative Learning (MACL) framework to address cross-modal alignment collapse in vision-language models when handling out-of-distributio…
Learning Dynamic Scene Reconstruction with Sinusoidal Geometric Priors
Tian Guo, Hui Yuan, Philip Xu +1
We propose SirenPose, a novel loss function that combines the periodic activation properties of sinusoidal representation networks with geometric priors derived from keypoint struc…
A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
Philip Xu
We introduce Uni4D, a unified framework for large scale open vocabulary 3D retrieval and controlled 4D generation based on structured three level alignment across text, 3D models,…
A Simple yet Effective Test-Time Adaptation for Zero-Shot Monocular Metric Depth Estimation
Rémi Marsal, Alexandre Chapoutot, Philippe Xu +1
The recent development of \emph{foundation models} for monocular depth estimation such as Depth Anything paved the way to zero-shot monocular depth estimation. Since it returns an…
Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?
Ziqi Ma, Sao Mai Nguyen, Philippe Xu
Emergent symbolic representations are critical for enabling developmental learning agents to plan and generalize across tasks. In this work, we investigate whether large language m…