10 papers
TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction
Nikos Athanasiou, Ilya A. Petrov, Angela Yao +8
Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal reasoning pipelines. However, r…
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
Qinchuan Cheng, Zhantao Gong, Pengzhan Sun +3
Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmark…
PRISM: : Planning and Reasoning with Intent in Simulated Embodied Environments
Yunn Kang Lim, Pengzhan Sun, Ziyi Bai +4
When an LLM-based embodied agent fails at a household task, the culprit could be misidentified objects, forgotten sub-goals, or poor action sequencing -- yet existing benchmarks re…
LightAVSeg: Lightweight Audio-Visual Segmentation
Qing Zhong, Guodong Ding, Lingqiao Liu +3
Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic…
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
Guodong Ding, Angela Yao
This paper tackles compositional personalization of vision-language models (VLMs). In this problem, multiple user-defined concepts must be recognized or described jointly at test t…
DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving
Xiaolu Liu, Yicong Li, Song Wang +3
Recently, world models have been incorporated into the autonomous driving systems to improve the planning reliability. Existing approaches typically predict future states through a…