papers

Publications (64)

cs.HC2022

Social Distancing Alert with Smartwatches

Xin Wang, Xilei Wu, Huina Meng +4

Social distancing is an efficient public health practice during the COVID-19 pandemic. However, people would violate the social distancing practice unconsciously when they conduct…

cs.CV2026

MobiDiary: Autoregressive Action Captioning with Wearable Devices and Wireless Signals

Fei Deng, Yinghui He, Chuntong Chu +4

Human Activity Recognition (HAR) in smart homes is critical for health monitoring and assistive living. While vision-based systems are common, they face privacy concerns and enviro…

cs.CV2023

Self-Supervised Image Representation Learning: Transcending Masking with Paired Image Overlay

Yinheng Li, Han Ding, Shaofei Wang

Self-supervised learning has become a popular approach in recent years for its ability to learn meaningful representations without the need for data annotation. This paper proposes…

cs.CV2026

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hengyi Xie, Chenfei Yao, Xianjin Wu +7

TurboVLA is a vision-language-action model that directly maps visual observations and language instructions to robot actions, achieving real-time performance (32 Hz) on an RTX 4090…

#vision-language models#robotic manipulation#real-time inference#lightweight architecture
cs.RO2026

DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference

Yuquan Li, Lianjie Ma, Han Ding +1

Vision-Language-Action (VLA) models enable generalist robotic manipulation but suffer from high inference latency. This bottleneck stems from the massive number of visual tokens pr…

cs.CV2026

Cross-domain EEG-based Emotion Recognition with Contrastive Learning

Rui Yan, Yibo Li, Han Ding +1

Electroencephalogram (EEG)-based emotion recognition is vital for affective computing but faces challenges in feature utilization and cross-domain generalization. This work introdu…

cs.CV2024

Diffusion Models For Multi-Modal Generative Modeling

Changyou Chen, Han Ding, Bunyamin Sisman +5

Diffusion-based generative modeling has been achieving state-of-the-art results on various generation tasks. Most diffusion models, however, are limited to a single-generation mode…

cs.CV2025

One Snapshot is All You Need: A Generalized Method for mmWave Signal Generation

Teng Huang, Han Ding, Wenxin Sun +6

Wireless sensing systems, particularly those using mmWave technology, offer distinct advantages over traditional vision-based approaches, such as enhanced privacy and effectiveness…

cs.RO2026

Optimal Uncertainty-Aware Calibration for the AX=YB Problem

Yanjia Chen, Xiangfei Li, Huan Zhao +4

This article proposes a general optimization framework for solving hand-eye calibration problem. Unlike traditional methods, an iterative algorithm based on Lie algebra that achiev…

cs.CL2025

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

MiniMax, :, Aili Chen +125

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combin…

cs.MA2026

Game-Theoretic Lens on LLM-based Multi-Agent Systems

Jianing Hao, Han Ding, Yuanjian Xu +5

Large language models (LLMs) have demonstrated strong reasoning, planning, and communication abilities, enabling them to operate as autonomous agents in open environments. While si…

cs.CL2025

A Systematic Survey of Automatic Prompt Optimization Techniques

Kiran Ramnath, Kang Zhou, Sheng Guan +18

Since the advent of large language models (LLMs), prompt engineering has been a crucial step for eliciting desired responses for various Natural Language Processing (NLP) tasks. Ho…

cs.RO2025

MonoSLAM: Robust Monocular SLAM with Global Structure Optimization

Bingzheng Jiang, Jiayuan Wang, Han Ding +1

This paper presents a robust monocular visual SLAM system that simultaneously utilizes point, line, and vanishing point features for accurate camera pose estimation and mapping. To…

stat.ML2019

Machine Discovery of Partial Differential Equations from Spatiotemporal Data

Ye Yuan, Junlin Li, Liang Li +9

The study presents a general framework for discovering underlying Partial Differential Equations (PDEs) using measured spatiotemporal data. The method, called Sparse Spatiotemporal…

cs.LG2026

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Jiacheng Chen, Xinyu Zhang, Shunkai Zhang +20

We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabili…

cs.HC2025

Active Domain Adaptation for mmWave-based HAR via Renyi Entropy-based Uncertainty Estimation

Mingzhi Lin, Teng Huang, Han Ding +4

Human Activity Recognition (HAR) using mmWave radar provides a non-invasive alternative to traditional sensor-based methods but suffers from domain shift, where model performance d…

cs.RO2026

Anatomical Prior-Driven Framework for Autonomous Robotic Cardiac Ultrasound Standard View Acquisition

Zhiyan Cao, Zhengxi Wu, Yiwei Wang +5

Cardiac ultrasound diagnosis is critical for cardiovascular disease assessment, but acquiring standard views remains highly operator-dependent. Existing medical segmentation models…

cs.GR2025

Conformal Slit Mapping Based Spiral Tool Trajectory Planning for Ball-end Milling on Complex Freeform Surfaces

Changqing Shen, BingZhou Xu, Xiaojian Zhang +2

This study presents a spiral-based complete coverage strategy for ball-end milling on freeform surfaces, utilizing conformal slit mapping to generate milling trajectories that are…

cs.SD2025

L3AC: Towards a Lightweight and Lossless Audio Codec

Linwei Zhai, Han Ding, Cui Zhao +4

Neural audio codecs have recently gained traction for their ability to compress high-fidelity audio and provide discrete tokens for generative modeling. However, leading approaches…

cs.LG2021

DeceFL: A Principled Decentralized Federated Learning Framework

Ye Yuan, Jun Liu, Dou Jin +12

Traditional machine learning relies on a centralized data pipeline, i.e., data are provided to a central server for model training. In many applications, however, data are inherent…

cs.NI2023

CRC-based Reliable WiFi Backscatter Communication for Supply Chain Management

Yun-Hao Liu, Tao Liu, Yimeng Huang +3

Supply chain management is aimed to keep going long-term performance of the supply chain and minimize the costs. Backscatter technology provides a more efficient way of being able…

cs.LG2026

VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation

Linwei Zhai, Han Ding, Mingzhi Lin +5

Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental to modern generative modeling, yet they often suffer from training instability and "codebook collapse" due to th…

cs.CV2025

mmEgoHand: Egocentric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMU

Yizhe Lv, Tingting Zhang, Zhijian Wang +4

Recent advancements in millimeter-wave (mmWave) radar have demonstrated its potential for human action recognition and pose estimation, offering privacy-preserving advantages over…

cs.RO2026

A Gait Driven Reinforcement Learning Framework for Humanoid Robots

Bolin Li, Yuzhi Jiang, Linwei Sun +3

This paper presents a real-time gait driven training framework for humanoid robots. First, we introduce a novel gait planner that incorporates dynamics to design the desired joint…

cs.RO2026

High-Load-Density Electro-Permanent Magnetic Foot with Controllable Adhesion for Quadruped Wall-Climbing Robots

An Li, Bo Tao, I-Ming Chen +1

To enable reliable climbing locomotion of quadruped robots on ferromagnetic surfaces, this paper presents a high-load-density electro-permanent magnetic foot with controllable adhe…

q-fin.TR2026

Large Language Model Agent in Financial Trading: A Survey

Han Ding, Yinheng Li, Junhao Wang +3

Trading is a highly competitive task that requires a combination of strategy, knowledge, and psychological fortitude. With the recent success of large language models(LLMs), it is…

cs.CV2025

HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving

R. D. Lin, Pengcheng Weng, Yinqiao Wang +3

LiDAR point cloud semantic segmentation plays a crucial role in autonomous driving. In recent years, semi-supervised methods have gained popularity due to their significant reducti…

cs.CV2025

XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses

Bo Lan, Pei Li, Jiaxi Yin +5

Human Action Recognition (HAR) plays a crucial role in applications such as health monitoring, smart home automation, and human-computer interaction. While HAR has been extensively…

eess.SY2018

Data-driven Discovery of Cyber-Physical Systems

Ye Yuan, Xiuchuan Tang, Wei Pan +5

Cyber-physical systems (CPSs) embed software into the physical world. They appear in a wide range of applications such as smart grids, robotics, intelligent manufacture and medical…

q-fin.GN2024

Large Language Models in Finance: A Survey

Yinheng Li, Shaofei Wang, Han Ding +1

Recent advances in large language models (LLMs) have opened new possibilities for artificial intelligence applications in finance. In this paper, we provide a practical survey focu…

cs.CV2026

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

Jiayu Gu, Yiwei Wang, Jie Zhang +14

Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to obtain because surgical int…

cs.RO2026

Stage-Aware and Roughness-Constrained Diffusion Policy for Multi-Stage Robotic Polishing

Shuai Ke, Jiexin Zhang, Huan Zhao +7

Polishing is a critical finishing process in high-end manufacturing fields such as aerospace, where surface quality directly affects the service performance and reliability of comp…

cs.RO2025

Leveraging Surgical Activity Grammar for Primary Intention Prediction in Laparoscopy Procedures

Jie Zhang, Song Zhou, Yiwei Wang +4

Surgical procedures are inherently complex and dynamic, with intricate dependencies and various execution paths. Accurate identification of the intentions behind critical actions,…

cs.RO2024

Spiral Complete Coverage Path Planning Based on Conformal Slit Mapping in Multi-connected Domains

Changqing Shen, Sihao Mao, Bingzhou Xu +4

The generation of smoother and shorter spiral complete coverage paths in multi-connected domains is a crucial research topic in path planning for robotic cavity machining and other…

cs.CV2025

Talk is Not Always Cheap: Promoting Wireless Sensing Models with Text Prompts

Zhenkui Yang, Zeyi Huang, Ge Wang +3

Wireless signal-based human sensing technologies, such as WiFi, millimeter-wave (mmWave) radar, and Radio Frequency Identification (RFID), enable the detection and interpretation o…

cs.RO2021

An Active Sense and Avoid System for Flying Robots in Dynamic Environments

Gang Chen, Wei Dong, Xinjun Sheng +2

This paper investigates a novel active-sensing-based obstacle avoidance paradigm for flying robots in dynamic environments. Instead of fusing multiple sensors to enlarge the field…

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

MiniMax, :, Aili Chen +219

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

cs.RO2024

Generation of Conservative Dynamical Systems Based on Stiffness Encoding

Tengyu Hou, Hanming Bai, Ye Ding +1

Dynamical systems (DSs) provide a framework for high flexibility, robustness, and control reliability and are widely used in motion planning and physical human-robot interaction. T…

cs.RO2026

Surface Constraint Policy for Learning Surface-Constrained and Dynamically Feasible Robot Skills

Shuai Ke, Jiexin Zhang, Huan Zhao +4

Diffusion-based imitation learning methods have driven rapid progress in robot dexterous manipulation tasks. However, they have limitations when applied to tasks that involve compl…

cs.RO2026

Learning-Based Adaptive Control for Surgical Robotic Exposure Task on Deformable Tissues

Jiayi Liu, Kaiqi Wei, Yiwei Wang +2

In various surgical procedures, regions of interest (ROIs) such as organs or lesions are often occluded by overlying tissues, requiring surgeons to achieve adequate exposure for pr…

cs.LG2022

A deep learning-based remaining useful life prediction approach for bearings

Cheng Cheng, Guijun Ma, Yong Zhang +4

In industrial applications, nearly half the failures of motors are caused by the degradation of rolling element bearings (REBs). Therefore, accurately estimating the remaining usef…

cs.RO2019

Learning to Navigate from Simulation via Spatial and Semantic Information Synthesis with Noise Model Embedding

Gang Chen, Hongzhe Yu, Wei Dong +3

While training an end-to-end navigation network in the real world is usually of high cost, simulation provides a safe and cheap environment in this training stage. However, trainin…

cs.RO2025

Robotic Grinding Skills Learning Based on Geodesic Length Dynamic Motion Primitives

Shuai Ke, Huan Zhao, Xiangfei Li +3

Learning grinding skills from human craftsmen via imitation learning has become a key research topic in robotic machining. Due to their strong generalization and robustness to exte…

cs.AI2026

CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection

Linbo Liu, Guande Wu, Han Ding +7

Large language model agents rely on effective model context to obtain task-relevant information for decision-making. Many existing context engineering approaches primarily rely on…

cs.CV2026

A Survey on Wi-Fi Sensing Generalizability: Taxonomy, Techniques, Datasets, and Future Research Prospects

Fei Wang, Tingting Zhang, Wei Xi +8

Wi-Fi sensing has emerged as a powerful non-intrusive technology for recognizing human activities, monitoring vital signs, and enabling context-aware applications using commercial…

cs.RO2026

Omnidirectional Humanoid Locomotion on Stairs via Unsafe Stepping Penalty and Sparse LiDAR Elevation Mapping

Yuzhi Jiang, Yujun Liang, Junhao Li +2

Humanoid robots, characterized by numerous degrees of freedom and a high center of gravity, are inherently unstable. Safe omnidirectional locomotion on stairs requires both omnidir…

cs.AI2025

SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond

Junteng Liu, Yuanxiang Fan, Zhuo Jiang +12

Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). Whi…

cs.CV2026

WS-IMUBench: Can Weakly Supervised Methods from Audio, Image, and Video Be Adapted for IMU-based Temporal Action Localization?

Pei Li, Jiaxi Yin, Lei Ouyang +4

IMU-based Human Activity Recognition (HAR) has enabled a wide range of ubiquitous computing applications, yet its dominant clip classification paradigm cannot capture the rich temp…

cs.LG2022

A General End-to-end Diagnosis Framework for Manufacturing Systems

Ye Yuan, Guijun Ma, Cheng Cheng +4

The manufacturing sector is envisioned to be heavily influenced by artificial intelligence-based technologies with the extraordinary increases in computational power and data volum…

cs.RO2026

MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks

Yaohang Xu, Lianjie Ma, Gewei Zuo +3

Reinforcement Learning (RL) and Imitation Learning (IL) are the standard frameworks for policy acquisition in manipulation. While IL offers efficient policy derivation, it suffers…

cs.RO2026

A Three-Level Whole-Body Disturbance Rejection Control Framework for Dynamic Motions in Legged Robots

Bolin Li, Gewei Zuo, Zhixiang Wang +3

This paper presents a control framework designed to enhance the stability and robustness of legged robots in the presence of uncertainties, including model uncertainties, external…

cs.CV2025

What's on Your Plate? Inferring Chinese Cuisine Intake from Wearable IMUs

Jiaxi Yin, Pengcheng Wang, Han Ding +1

Accurate food intake detection is vital for dietary monitoring and chronic disease prevention. Traditional self-report methods are prone to recall bias, while camera-based approach…

eess.SP2025

You Can Wash Hands Better: Accurate Daily Handwashing Assessment with a Smartwatch

Fei Wang, Tingting Zhang, Xilei Wu +6

Hand hygiene is among the most effective daily practices for preventing infectious diseases such as influenza, malaria, and skin infections. While professional guidelines emphasize…

cs.RO2026

AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory

Lianjie Ma, Yuquan Li, Bingzheng Jiang +3

Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational cost often prohibits deployment on edge…

cs.RO2025

A Whole-Body Disturbance Rejection Control Framework for Dynamic Motions in Legged Robots

Bolin Li, Wentao Zhang, Xuecong Huang +2

This letter presents a control framework for legged robots that enables self-perception and resistance to external disturbances and model uncertainties. First, a novel disturbance…

cs.IR2026

What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation

Kainan Shi, Peilin Zhou, Ge Wang +2

Using Large Language Models (LLMs) to generate semantic features has been demonstrated as a powerful paradigm for enhancing Sequential Recommender Systems (SRS). This typically inv…

cs.CL2023

Cross-corpus Readability Compatibility Assessment for English Texts

Zhenzhen Li, Han Ding, Shaohong Zhang

Text readability assessment has gained significant attention from researchers in various domains. However, the lack of exploration into corpus compatibility poses a challenge as di…

cs.RO2025

Disturbance Estimation of Legged Robots: Predefined Convergence via Dynamic Gains

Bolin Li, Peiyuan Cai, Gewei Zuo +2

In this study, we address the challenge of disturbance estimation in legged robots by introducing a novel continuous-time online feedback-based disturbance observer that leverages…

cs.AI2025

Bridging Social Psychology and LLM Reasoning: Conflict-Aware Meta-Review Generation via Cognitive Alignment

Wei Chen, Han Ding, Meng Yuan +3

The rapid growth of scholarly submissions has overwhelmed traditional peer review systems, driving the need for intelligent automation to preserve scientific rigor. While large lan…

cs.LG2024

CCS: Continuous Learning for Customized Incremental Wireless Sensing Services

Qunhang Fu, Fei Wang, Mengdie Zhu +3

Wireless sensing has made significant progress in tasks ranging from action recognition, vital sign estimation, pose estimation, etc. After over a decade of work, wireless sensing…

cs.SD2025

We Can Hear You with mmWave Radar! An End-to-End Eavesdropping System

Dachao Han, Teng Huang, Han Ding +4

With the rise of voice-enabled technologies, loudspeaker playback has become widespread, posing increasing risks to speech privacy. Traditional eavesdropping methods often require…

cs.CV2024

Data Processing Techniques for Modern Multimodal Models

Yinheng Li, Han Ding, Hang Chen

Data processing plays an significant role in current multimodal model training. In this paper. we provide an comprehensive review of common data processing techniques used in moder…

cs.HC2022

Mask Wearing Status Estimation with Smartwatches

Huina Meng, Xilei Wu, Xin Wang +4

We present MaskReminder, an automatic mask-wearing status estimation system based on smartwatches, to remind users who may be exposed to the COVID-19 virus transmission scenarios,…

cs.CV2024

Bringing Multimodality to Amazon Visual Search System

Xinliang Zhu, Michael Huang, Han Ding +10

Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns betw…