11 citations · 20 across the 25 of their papers we have counts for
10 papers · 1 filter
Hierarchically Robust Zero-shot Vision-language Models
Junhao Dong, Yifei Zhang, Hao Zhu +2
Vision-Language Models (VLMs) can perform zero-shot classification but are susceptible to adversarial attacks. While robust fine-tuning improves their robustness, existing approach…
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
Heng Zhao, Yew-Soon Ong, Joey Tianyi Zhou
Spatio-Temporal Video Grounding (STVG) aims to retrieve the spatio-temporal tube of a target object or person in a video given a text query. Most existing approaches perform frame-…
NeuSpring: Neural Spring Fields for Reconstruction and Simulation of Deformable Objects from Videos
Qingshan Xu, Jiao Liu, Shangshu Yu +6
In this paper, we aim to create physical digital twins of deformable objects under interaction. Existing methods focus more on the physical learning of current state modeling, but…
Lightweight and Accurate Multi-View Stereo with Confidence-Aware Diffusion Model
Fangjinhua Wang, Qingshan Xu, Yew-Soon Ong +1
To reconstruct the 3D geometry from calibrated images, learning-based multi-view stereo (MVS) methods typically perform multi-view depth estimation and then fuse depth maps into a…
LLM-to-Phy3D: Physically Conform Online 3D Object Generation with LLMs
Melvin Wong, Yueming Lyu, Thiago Rios +2
The emergence of generative artificial intelligence (GenAI) and large language models (LLMs) has revolutionized the landscape of digital content creation in different modalities. H…
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
Yinjie Zhao, Heng Zhao, Bihan Wen +2
With the rapid development of vision tasks and the scaling on datasets and models, redundancy reduction in vision datasets has become a key area of research. To address this issue,…