6 papers
SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
Sen Fang, Hongbin Zhong, Yanxin Zhang +1
Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such r…
Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models
Shuibai Zhang, Caspian Zhuang, Chihan Cui +8
Diffusion language models (DLMs) enable parallel, non-autoregressive text generation, yet existing DLM mixture-of-experts (MoE) models inherit token-choice (TC) routing from autore…
HabitatAgent: An End-to-End Multi-Agent System for Housing Consultation
Hongyang Yang, Yanxin Zhang, Yang She +5
Housing selection is a high-stakes and largely irreversible decision problem. We study housing consultation as a decision-support interface for housing selection. Existing housing…
RAC: Rectified Flow Auto Coder
Sen Fang, Yalin Feng, Yanxin Zhang +1
In this paper, we propose a Rectified Flow Auto Coder (RAC) inspired by Rectified Flow to replace the traditional VAE: 1. It achieves multi-step decoding by applying the decoder to…
LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
Zeyi Kang, Liang He, Yanxin Zhang +2
Multimodal semantic learning plays a critical role in embodied intelligence, especially when robots perceive their surroundings, understand human instructions, and make intelligent…
M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
Yanxin Zhang, Liang He, Zeyi Kang +2
In recent years, multimodal learning has become essential in robotic vision and information fusion, especially for understanding human behavior in complex environments. However, cu…