papers

Publications (54)

cs.LG2025

Learning Robust Spectral Dynamics for Temporal Domain Generalization

En Yu, Jie Lu, Xiaoyu Yang +2

Modern machine learning models struggle to maintain performance in dynamic environments where temporal distribution shifts, \emph{i.e., concept drift}, are prevalent. Temporal Doma…

cs.AI2025

Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction

Bao Shu, Yan Cai, Jianjian Sun +11

Developing robust world model reasoning is crucial for large language model (LLM) agents to plan and interact in complex environments. While multi-turn interaction offers a superio…

cs.AI2025

PerPO: Perceptual Preference Optimization via Discriminative Rewarding

Zining Zhu, Liang Zhao, Kangheng Lin +7

This paper presents Perceptual Preference Optimization (PerPO), a perception alignment method aimed at addressing the visual discrimination challenges in generative pre-trained mul…

cs.CV2024

Small Language Model Meets with Reinforced Vision Vocabulary

Haoran Wei, Lingyu Kong, Jinyue Chen +6

Playing Large Vision Language Models (LVLMs) in 2023 is trendy among the AI community. However, the relatively large number of parameters (more than 7B) of popular LVLMs makes it d…

cs.LG2026

Bandwidth-constrained Variational Message Encoding for Cooperative Multi-agent Reinforcement Learning

Wei Duan, Jie Lu, En Yu +1

Graph-based multi-agent reinforcement learning (MARL) enables coordinated behavior under partial observability by modeling agents as nodes and communication links as edges. While r…

cs.CV2026

RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking

Yanqiu Yu, Zhifan Jin, Sijia Chen +4

Referring Multi-Object Tracking has attracted increasing attention due to its human-friendly interactive characteristics, yet it exhibits limitations in low-visibility conditions,…

cs.CV2025

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

Yana Wei, Liang Zhao, Jianjian Sun +15

The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates…

cs.AI2026

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation

Changhua Xu, En Yu, Junyu Xuan +1

Vision--Language--Action (VLA) models bridge multimodal reasoning with physical control, but adapting them to new tasks with scarce demonstrations remains unreliable. While fine-tu…

cs.LG2026

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning

Xiaoyu Yang, En Yu, Wei Duan +1

Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with complex human values and domain-sp…

cs.LG2026

Autonomous Drift Learning in Data Streams: A Unified Perspective

Xiaoyu Yang, En Yu, Jie Lu

In the pursuit of autonomous learning systems, the foundational assumption of stationarity, the premise that data distributions and model behaviors remain constant, is fundamentall…

cs.CV2025

Unhackable Temporal Rewarding for Scalable Video MLLMs

En Yu, Kangheng Lin, Liang Zhao +8

In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the "anti-scaling law", where more data and larger models lead to worse performance. Th…

cs.IR2020

Fusion-supervised Deep Cross-modal Hashing

Li Wang, Lei Zhu, En Yu +2

Deep hashing has recently received attention in cross-modal retrieval for its impressive advantages. However, existing hashing methods for cross-modal retrieval cannot fully captur…

cs.CV2026

STEP3-VL-10B Technical Report

Ailin Huang, Chengyuan Yao, Chunrui Han +90

We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-…

cs.LG2025

Drift-aware Collaborative Assistance Mixture of Experts for Heterogeneous Multistream Learning

En Yu, Jie Lu, Kun Wang +2

Learning from multiple data streams in real-world scenarios is fundamentally challenging due to intrinsic heterogeneity and unpredictable concept drifts. Existing methods typically…

cs.CL2025

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios

Ruiwen Zhou, Wenyue Hua, Liangming Pan +4

This paper introduces RuleArena, a novel and challenging benchmark designed to evaluate the ability of large language models (LLMs) to follow complex, real-world rules in reasoning…

cs.CV2024

Cross-View Referring Multi-Object Tracking

Sijia Chen, En Yu, Wenbing Tao

Referring Multi-Object Tracking (RMOT) is an important topic in the current tracking field. Its task form is to guide the tracker to track objects that match the language descripti…

cs.CV2025

Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion

Enyu Liu, En Yu, Sijia Chen +1

3D Semantic Scene Completion (SSC) has gained increasing attention due to its pivotal role in 3D perception. Recent advancements have primarily focused on refining voxel-level feat…

cs.CV2026

ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking

Sijia Chen, Zihan Zhou, Yanqiu Yu +2

Multi-Object Tracking (MOT) is a fundamental task in computer vision, aiming to track targets across video frames. Existing MOT methods perform well in general visual scenes, but f…

cs.CV2025

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

NextStep Team, Chunrui Han, Guopeng Li +47

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ ve…

cs.CV2022

Generalizing Multiple Object Tracking to Unseen Domains by Introducing Natural Language Representation

En Yu, Songtao Liu, Zhuoling Li +4

Although existing multi-object tracking (MOT) algorithms have obtained competitive performance on various benchmarks, almost all of them train and validate models on the same domai…

cs.CV2025

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

En Yu, Kangheng Lin, Liang Zhao +11

Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, ou…

cs.CV2025

OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer

Jinyang Li, En Yu, Sijia Chen +1

Open-vocabulary multiple object tracking aims to generalize trackers to unseen categories during training, enabling their application across a variety of real-world scenarios. Howe…

cs.CV2026

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

Fanfu Xue, En Yu, Yantian Shen +5

UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized a…

cs.LG2025

Autonomous Concept Drift Threshold Determination

Pengqian Lu, Jie Lu, Anjin Liu +2

Existing drift detection methods focus on designing sensitive test statistics. They treat the detection threshold as a fixed hyperparameter, set once to balance false alarms and la…

cs.LG2026

Generalized Incremental Learning under Concept Drift across Evolving Data Streams

En Yu, Jie Lu, Guangquan Zhang

Real-world data streams exhibit inherent non-stationarity characterized by concept drift, posing significant challenges for adaptive learning systems. While existing methods addres…

cs.CV2026

DRMOT: A Dataset and Framework for RGBD Referring Multi-Object Tracking

Sijia Chen, Lijuan Ma, Yanqiu Yu +3

Referring Multi-Object Tracking (RMOT) aims to track specific targets based on language descriptions and is vital for interactive AI systems such as robotics and autonomous driving…

cs.LG2025

Walking the Tightrope: Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning

Xiaoyu Yang, Jie Lu, En Yu

This paper uncovers a critical yet overlooked phenomenon in multi-modal large language models (MLLMs): detrimental concept drift within chain-of-thought (CoT) reasoning during non-…

cs.CV2025

InstaFace: Identity-Preserving Facial Editing with Single Image Inference

MD Wahiduzzaman Khan, Mingshan Jia, Xiaolin Zhang +3

Facial appearance editing is crucial for digital avatars, AR/VR, and personalized content creation, driving realistic user experiences. However, preserving identity with generative…

cs.CV2022

Quality Matters: Embracing Quality Clues for Robust 3D Multi-Object Tracking

Jinrong Yang, En Yu, Zeming Li +2

3D Multi-Object Tracking (MOT) has achieved tremendous achievement thanks to the rapid development of 3D object detection and 2D MOT. Recent advanced works generally employ a serie…

cs.CV2021

RelationTrack: Relation-aware Multiple Object Tracking with Decoupled Representation

En Yu, Zhuoling Li, Shoudong Han +1

Existing online multiple object tracking (MOT) algorithms often consist of two subtasks, detection and re-identification (ReID). In order to enhance the inference speed and reduce…

cs.RO2026

DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI

En Yu, Haoran Lv, Jianjian Sun +46

Moving beyond the traditional paradigm of adapting internet-pretrained models to physical tasks, we present DM0, an Embodied-Native Vision-Language-Action (VLA) framework designed…

cs.CV2026

Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars

MD Wahiduzzaman Khan, Mingshan Jia, Xiaolin Zhang +2

The paper introduces SpiD, a framework that creates photorealistic, animatable head avatars from a single image using a dual-axis disentanglement of Gaussian representations, enabl…

#head avatars#gaussian splatting#real-time rendering#facial animation
cs.CV2024

Delving into the Trajectory Long-tail Distribution for Muti-object Tracking

Sijia Chen, En Yu, Jinyang Li +1

Multiple Object Tracking (MOT) is a critical area within computer vision, with a broad spectrum of practical implementations. Current research has primarily focused on the developm…

cs.CV2022

Delving into the Pre-training Paradigm of Monocular 3D Object Detection

Zhuoling Li, Chuanrui Zhang, En Yu +1

The labels of monocular 3D object detection (M3OD) are expensive to obtain. Meanwhile, there usually exists numerous unlabeled data in practical applications, and pre-training is a…

cs.LG2024

Online Boosting Adaptive Learning under Concept Drift for Multistream Classification

En Yu, Jie Lu, Bin Zhang +1

Multistream classification poses significant challenges due to the necessity for rapid adaptation in dynamic streaming processes with concept drift. Despite the growing research ou…

cs.CV2020

MAT: Motion-Aware Multi-Object Tracking

Shoudong Han, Piao Huang, Hongwei Wang +4

Modern multi-object tracking (MOT) systems usually model the trajectories by associating per-frame detections. However, when camera motion, fast motion, and occlusion challenges oc…

cs.CV2020

Refinements in Motion and Appearance for Online Multi-Object Tracking

Piao Huang, Shoudong Han, Jun Zhao +4

Modern multi-object tracking (MOT) system usually involves separated modules, such as motion model for location and appearance model for data association. However, the compatible p…

cs.CV2025

Perception in Reflection

Yana Wei, Liang Zhao, Kangheng Lin +8

We present a perception in reflection paradigm designed to transcend the limitations of current large vision-language models (LVLMs), which are expected yet often fail to achieve p…

cs.CV2024

Merlin:Empowering Multimodal LLMs with Foresight Minds

En Yu, Liang Zhao, Yana Wei +8

Humans possess the remarkable ability to foresee the future to a certain extent based on present observations, a skill we term as foresight minds. However, this capability remains…

cs.CV2023

GroupLane: End-to-End 3D Lane Detection with Channel-wise Grouping

Zhuoling Li, Chunrui Han, Zheng Ge +5

Efficiency is quite important for 3D lane detection due to practical deployment demand. In this work, we propose a simple, fast, and end-to-end detector that still maintains high d…

cs.CV2025

Adapting Multi-modal Large Language Model to Concept Drift From Pre-training Onwards

Xiaoyu Yang, Jie Lu, En Yu

Multi-modal Large Language Models (MLLMs) frequently face challenges from concept drift when dealing with real-world streaming data, wherein distributions change unpredictably. Thi…

cs.CV2026

Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments

Xiaoyu Yang, En Yu, Wei Duan +1

This paper identifies a critical yet underexplored challenge in reasoning alignment from multiple multi-modal large language models (MLLMs): In non-stationary environments, the div…

cs.CV2023

MOTRv3: Release-Fetch Supervision for End-to-End Multi-Object Tracking

En Yu, Tiancai Wang, Zhuoling Li +3

Although end-to-end multi-object trackers like MOTR enjoy the merits of simplicity, they suffer from the conflict between detection and association seriously, resulting in unsatisf…

cs.CV2022

Towards Discriminative Representation: Multi-view Trajectory Contrastive Learning for Online Multi-object Tracking

En Yu, Zhuoling Li, Shoudong Han

Discriminative representation is crucial for the association step in multi-object tracking. Recent work mainly utilizes features in single or neighboring frames for constructing me…

cs.LG2025

Resilient Contrastive Pre-training under Non-Stationary Drift

Xiaoyu Yang, Jie Lu, En Yu +1

The remarkable success of large-scale contrastive pre-training has been largely driven by by vast yet static datasets. However, as the scaling paradigm evolves, this paradigm encou…

cs.CV2026

TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

MD Wahiduzzaman Khan, Mingshan Jia, Xiaolin Zhang +2

The paper presents a framework that adds realistic, cross‑identity tongue motion to face reenactment by automatically training a tongue segmentation model and using a spatially con…

#face reenactment#tongue synthesis#diffusion models#segmentation
cs.RO2026

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation

Fanfu Xue, En Yu, Bohang Liu +4

UAV see-and-reach navigation requires an aerial agent to approach a language-specified target visible in its initial view and stop reliably near it. Existing methods typically map…

cs.AI2026

MentalThink: Shaping Thoughts in Mental SVG World

Kangheng Lin, Jisheng Yin, Dingming Li +11

We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink…

cs.AI2026

Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning

Wei Duan, Junyu Xuan, En Yu +2

Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners lack a theoretically grounded mechanism t…

cs.CV2026

ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking

Sijia Chen, Yanqiu Yu, En Yu +1

Referring Multi-Object Tracking (RMOT) aims to track targets specified by language instructions. However, existing RMOT paradigms heavily rely on explicit visual-textual matching a…

cs.LG2025

Multimodal Inverse Attention Network with Intrinsic Discriminant Feature Exploitation for Fake News Detection

Tianlin Zhang, En Yu, Yi Shao +1

Multimodal fake news detection has garnered significant attention due to its profound implications for social security. While existing approaches have contributed to understanding…

cs.CV2026

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Yana Wei, Hongbo Peng, Yanlin Lai +14

We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from h…

cs.CV2023

Implicit and Efficient Point Cloud Completion for 3D Single Object Tracking

Pan Wang, Liangliang Ren, Shengkai Wu +4

The point cloud based 3D single object tracking has drawn increasing attention. Although many breakthroughs have been achieved, we also reveal two severe issues. By extensive analy…

cs.CL2023

ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Liang Zhao, En Yu, Zheng Ge +8

Human-AI interactivity is a critical aspect that reflects the usability of multimodal large language models (MLLMs). However, existing end-to-end MLLMs only allow users to interact…