papers

Publications (31)

cs.CV2025

Temporal Action Detection Model Compression by Progressive Block Drop

Xiaoyong Chen, Yong Guo, Jiaming Liang +3

Temporal action detection (TAD) aims to identify and localize action instances in untrimmed videos, which is essential for various video understanding tasks. However, recent improv…

cs.CV2025

Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models

Ying Peng, Hongsen Ye, Changxin Huang +3

Vision Transformers (ViTs) have achieved strong performance in video action recognition, but their high computational cost limits their practicality. Lightweight CNNs are more effi…

cs.RO2024

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution

Changxin Huang, Yanbin Chang, Junfan Lin +3

The ability to autonomously explore and resolve tasks with minimal human guidance is crucial for the self-development of embodied intelligence. Although reinforcement learning meth…

cs.CV2025

Towards Stable Cross-Domain Depression Recognition under Missing Modalities

Jiuyi Chen, Mingkui Tan, Haifeng Lu +4

Depression poses serious public health risks, including suicide, underscoring the urgency of timely and scalable screening. Multimodal automatic depression detection (ADD) offers a…

cs.CV2025

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

Runhao Zeng, Jiaqi Mao, Minghao Lai +5

Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streami…

eess.IV2020

A Thorough Comparison Study on Adversarial Attacks and Defenses for Common Thorax Disease Classification in Chest X-rays

Chendi Rao, Jiezhang Cao, Runhao Zeng +4

Recently, deep neural networks (DNNs) have made great progress on automated diagnosis with chest X-rays images. However, DNNs are vulnerable to adversarial examples, which may caus…

cs.LG2023

DCIR: Dynamic Consistency Intrinsic Reward for Multi-Agent Reinforcement Learning

Kunyang Lin, Yufeng Wang, Peihao Chen +4

Learning optimal behavior policy for each agent in multi-agent systems is an essential yet difficult problem. Despite fruitful progress in multi-agent reinforcement learning, the c…

cs.LG2026

DeRelayL: Sustainable Decentralized Relay Learning

Haihan Duan, Tengfei Ma, Yuyang Qin +4

In the era of big data, large-scale machine learning models have revolutionized various fields, driving significant advancements. However, large-scale model training demands high f…

cs.CV2026

Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models

Jingrui Zhang, Feng Liang, Yong Zhang +3

With the remarkable success of large language models (LLMs) in natural language understanding and generation, multimodal large language models (MLLMs) have rapidly advanced in thei…

cs.RO2025

Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators

Qiwei Liang, Boyang Cai, Rongyi He +5

Quadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, exis…

cs.CV2019

Graph Convolutional Networks for Temporal Action Localization

Runhao Zeng, Wenbing Huang, Mingkui Tan +4

Most state-of-the-art action localization systems process each action proposal individually, without explicitly exploiting their relations during learning. However, the relations b…

cs.HC2024

Understanding Emotional Body Expressions via Large Language Models

Haifeng Lu, Jiuyi Chen, Feng Liang +3

Emotion recognition based on body movements is vital in human-computer interaction. However, existing emotion recognition methods predominantly focus on enhancing classification ac…

cs.RO2024

Video2Reward: Generating Reward Function from Videos for Legged Robot Behavior Learning

Runhao Zeng, Dingjie Zhou, Qiwei Liang +6

Learning behavior in legged robots presents a significant challenge due to its inherent instability and complex constraints. Recent research has proposed the use of a large languag…

cs.CV2021

RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning

Peihao Chen, Deng Huang, Dongliang He +5

We study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such…

cs.CV2024

Benchmarking the Robustness of Temporal Action Detection Models Against Temporal Corruptions

Runhao Zeng, Xiaoyong Chen, Jiaming Liang +3

Temporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Although many methods have achieved promising results…

cs.CV2022

Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation

Peihao Chen, Dongyu Ji, Kunyang Lin +4

We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions oft…

cs.CV2024

Towards Long Video Understanding via Fine-detailed Video Story Generation

Zeng You, Zhiquan Wen, Yaofo Chen +4

Long video understanding has become a critical task in computer vision, driving advancements across numerous applications from surveillance to content retrieval. Existing video und…

cs.CV2021

Graph Convolutional Module for Temporal Action Localization in Videos

Runhao Zeng, Wenbing Huang, Mingkui Tan +4

Temporal action localization has long been researched in computer vision. Existing state-of-the-art action localization methods divide each video into multiple action units (i.e.,…

cs.LG2024

Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers

Qi Deng, Shuaicheng Niu, Ronghao Zhang +4

Test-time adaptation (TTA) aims to fine-tune a trained model online using unlabeled testing data to adapt to new environments or out-of-distribution data, demonstrating broad appli…

cs.CV2023

Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Peihao Chen, Xinyu Sun, Hongyan Zhi +5

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by langu…

cs.CV2026

AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries

Zhen Zhang, Yuhang Yang, Yunxiang Jiang +5

Existing affective understanding studies have mainly focused on recognizing emotions from images, audio signals, or pre-cliped video clips, where the affective evidence is already…

cs.AI2026

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

Changxin Huang, Lv Tang, Zhaohuan Zhan +5

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging.…

cs.CV2025

Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation

Runhao Zeng, Qi Deng, Ronghao Zhang +4

Test-time adaptation (TTA) aims to boost the generalization capability of a trained model by conducting self-/unsupervised learning during the testing phase. While most existing TT…

cs.CV2026

ZOTTA: Test-Time Adaptation with Gradient-Free Zeroth-Order Optimization

Ronghao Zhang, Shuaicheng Niu, Qi Deng +3

Test-time adaptation (TTA) aims to improve model robustness under distribution shifts by adapting to unlabeled test data, but most existing methods rely on backpropagation (BP), wh…

cs.CV2020

Dense Regression Network for Video Grounding

Runhao Zeng, Haoming Xu, Wenbing Huang +3

We address the problem of video grounding from natural language queries. The key challenge in this task is that one training video might only contain a few annotated starting/endin…

cs.LG2025

Nesterov-Accelerated Robust Federated Learning Over Byzantine Adversaries

Lihan Xu, Yanjie Dong, Gang Wang +3

We investigate robust federated learning, where a group of workers collaboratively train a shared model under the orchestration of a central server in the presence of Byzantine adv…

cs.CV2026

Opening the Black Box: Preliminary Insights into Affective Modeling in Multimodal Foundation Models

Zhen Zhang, Runhao Zeng, Sicheng Zhao +1

Understanding where and how emotions are represented in large-scale foundation models remains an open problem, particularly in multimodal affective settings. Despite the strong emp…

cs.CV2020

Location-aware Graph Convolutional Networks for Video Question Answering

Deng Huang, Peihao Chen, Runhao Zeng +3

We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art method…

cs.LG2019

Continual Reinforcement Learning with Diversity Exploration and Adversarial Self-Correction

Fengda Zhu, Xiaojun Chang, Runhao Zeng +1

Deep reinforcement learning has made significant progress in the field of continuous control, such as physical control and autonomous driving. However, it is challenging for a rein…

cs.LG2025

CO-PFL: Contribution-Oriented Personalized Federated Learning for Heterogeneous Networks

Ke Xing, Yanjie Dong, Xiaoyi Fan +4

Personalized federated learning (PFL) addresses a critical challenge of collaboratively training customized models for clients with heterogeneous and scarce local data. Conventiona…

cs.CV2025

Emotion Recognition from Skeleton Data: A Comprehensive Survey

Haifeng Lu, Jiuyi Chen, Zhen Zhang +3

Emotion recognition through body movements has emerged as a compelling and privacy-preserving alternative to traditional methods that rely on facial expressions or physiological si…