collaborators

6 papers

cs.CV2026

SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages

Sen Fang, Hongbin Zhong, Yanxin Zhang +1

Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such r…

cs.LG2026

Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models

Shuibai Zhang, Caspian Zhuang, Chihan Cui +8

Diffusion language models (DLMs) enable parallel, non-autoregressive text generation, yet existing DLM mixture-of-experts (MoE) models inherit token-choice (TC) routing from autore…

cs.LG2026

HabitatAgent: An End-to-End Multi-Agent System for Housing Consultation

Hongyang Yang, Yanxin Zhang, Yang She +5

Housing selection is a high-stakes and largely irreversible decision problem. We study housing consultation as a decision-support interface for housing selection. Existing housing…

cs.CV2026

RAC: Rectified Flow Auto Coder

Sen Fang, Yalin Feng, Yanxin Zhang +1

In this paper, we propose a Rectified Flow Auto Coder (RAC) inspired by Rectified Flow to replace the traditional VAE: 1. It achieves multi-step decoding by applying the decoder to…

cs.RO2025

LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA

Zeyi Kang, Liang He, Yanxin Zhang +2

Multimodal semantic learning plays a critical role in embodied intelligence, especially when robots perceive their surroundings, understand human instructions, and make intelligent…

cs.RO2025

M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer

Yanxin Zhang, Liang He, Zeyi Kang +2

In recent years, multimodal learning has become essential in robotic vision and information fusion, especially for understanding human behavior in complex environments. However, cu…