activity
20242026
collaborators

9 papers

cs.LG2026

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Zirui Cheng, Xun Xu, Tiankai Chen +7

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the…

cs.CV2026

Thinking Ahead: Foresight Intelligence in MLLMs and World Model

Zhantao Gong, Liaoyuan Fan, Qing Guo +3

In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet lar…

cs.RO2026

PRISM: : Planning and Reasoning with Intent in Simulated Embodied Environments

Yunn Kang Lim, Pengzhan Sun, Ziyi Bai +4

When an LLM-based embodied agent fails at a household task, the culprit could be misidentified objects, forgotten sub-goals, or poor action sequencing -- yet existing benchmarks re…

cs.CV2025

Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation

Tiankai Chen, Yushu Li, Adam Goodge +4

Out-of-distribution (OOD) detection in 3D point cloud data remains a challenge, particularly in applications where safe and robust perception is critical. While existing OOD detect…

cs.CV2025

SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation

Adam Goodge, Xun Xu, Bryan Hooi +4

As point cloud data increases in prevalence in a variety of applications, the ability to detect out-of-distribution (OOD) point cloud objects becomes critical for ensuring model sa…

cs.CV2025

Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion

Yifan Liu, Xun Xu, Shijie Li +2

Multi-camera systems provide richer contextual information for industrial anomaly detection. However, traditional methods process each view independently, disregarding the compleme…