activity
20242026
collaborators

5 papers

cs.CV2026

SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

Zhewen He, Junyi Hu, Haomian Huang +3

Sign language models are typically trained on datasets captured under constrained conditions, with limited viewpoint, background, and signer-identity diversity, leading to poor rob…

cs.CV2026

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

Jiwen Liu, Shujuan Li, Zhixue Fang +8

Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametr…

cs.CV2025

Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation

Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala +3

Instance Image-Goal Navigation (IIN) requires autonomous agents to identify and navigate to a target object or location depicted in a reference image captured from any viewpoint. W…

cs.RO2025

Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language Models

Congcong Wen, Yifan Liu, Geeta Chandra Raju Bethala +7

Robot navigation is crucial across various domains, yet traditional methods focus on efficiency and obstacle avoidance, often overlooking human behavior in shared spaces. With the…

cs.RO2024

Zero-shot Object Navigation with Vision-Language Models Reasoning

Congcong Wen, Yisiyuan Huang, Hao Huang +6

Object navigation is crucial for robots, but traditional methods require substantial training data and cannot be generalized to unknown environments. Zero-shot object navigation (Z…