7 papers
Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models
Xiatao Sun, Yuan Zhuang, Mateo Sanchez Lopez Negrete +9
Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet robust deployment acros…
Open-World Video Segmentation
Qing Su, Kaiyang Li, Yuan Zhuang +2
While video segmentation has advanced rapidly on short clips and closed-set benchmarks, open-world video segmentation remains largely unexplored. The challenge is twofold: (1) exis…
SEVO: Semantic-Enhanced Virtual Observation for Robust VLA Manipulation via Active Illumination and Data-Centric Collection
Tianchonghui Fang, Yuan Zhuang, Fei Miao
Vision-Language-Action (VLA) and imitation-learning policies trained via community toolchains on low-cost hardware frequently fail when deployed outside the training environment. E…
LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
Yuan Zhuang, Yi Shen, Yuexin Bian +4
Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to…
TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
Chenchen Lin, Sanbao Su, Rachel Luo +4
Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's rich hierarchy of visual cues under-utili…
Uncertainty Quantification for Collaborative Object Detection Under Adversarial Attacks
Huiqun Huang, Cong Chen, Jean-Philippe Monteuuis +2
Collaborative Object Detection (COD) and collaborative perception can integrate data or features from various entities, and improve object detection accuracy compared with individu…