8 papers
Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding
Junyi Hu, Tian Bai, Fengyi Wu +7
Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this w…
Learning Adaptive Safety Margins for Visual Navigation
Junyi Hu, Shuaihang Yuan, Geeta Chandra Raju Bethala +2
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause…
VTaMo: Video-Text Alignment Model for Sign Language Translation
Junyi Hu, Zhewen He, Haomian Huang +2
Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely o…
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
Zhewen He, Junyi Hu, Haomian Huang +3
Sign language models are typically trained on datasets captured under constrained conditions, with limited viewpoint, background, and signer-identity diversity, leading to poor rob…
LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
Junyi Hu, Qiji Zhou, Lei Zhang +1
Zero-shot recognition aims to classify an image by selecting the most compatible label description from a set of candidate classes without any task-specific supervision. In fine-gr…
PHCT: Plug-and-Play Hierarchical C2F Transformer for Multi-Scale Feature Fusion
Junyi Hu, Tian Bai, Fengyi Wu +2
Feature fusion plays a pivotal role in achieving high performance in vision models, yet existing attention-based fusion techniques often suffer from substantial computational overh…