collaborators

6 papers

cs.RO2026

GeoTLM: Geometry-aware Tactile-Language Models for Contact Motion Orientation Reasoning of Dynamic Objects

Qiutian Li, Zinan Liu, Lin Wang

Modern tactile-language models (TLMs) have shown potential for robot learning tasks, such as material and texture recognition. However, for contact-rich scenarios, these TLMs strug…

cs.RO2026

VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI

Zinan Liu, Eric Zheng, Soumyaratna Debnath +3

Spiking neural networks (SNNs) are brain-inspired, event-driven models that compute with sparse spikes, which enables highly efficient visual perception in resource-constrained emb…

cs.CV2026

Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection

Sa Zhu, Wanqian Zhang, Lin Wang +3

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effective alignment between action…

cs.CV2026

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

Ze Dong, Hao Shi, Zejia Gao +3

Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shifts such as viewpoint distorti…

cs.RO2026

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

Hao Shi, Ze Wang, Shangwei Guo +6

Robust 3D semantic occupancy is crucial for legged/humanoid robots, yet most semantic scene completion (SSC) systems target wheeled platforms with forward-facing sensors. We presen…

cs.CV2025

Event-guided 3D Gaussian Splatting for Dynamic Human and Scene Reconstruction

Xiaoting Yin, Hao Shi, Kailun Yang +4

Reconstructing dynamic humans together with static scenes from monocular videos remains difficult, especially under fast motion, where RGB frames suffer from motion blur. Event cam…