papers

Publications (83)

cs.CV2025

AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows

Zhenglin Zhou, Fan Ma, Chengzhuo Gui +4

Training-free 3D editing aims to modify 3D shapes based on human instructions without model finetuning. It plays a crucial role in 3D content creation. However, existing approaches…

cs.LG2026

Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL

Yixiao Zhou, Yang Li, Dongzhou Cheng +2

Reinforcement Learning from Verifiable Rewards (RLVR) trains large language models (LLMs) from sampled trajectories, making decoding strategy a core component of learning rather th…

cs.CV2025

TV-Dialogue: Crafting Theme-Aware Video Dialogues with Immersive Interaction

Sai Wang, Fan Ma, Xinyi Li +2

Recent advancements in LLMs have accelerated the development of dialogue generation across text and images, yet video-based dialogue generation remains underexplored and presents u…

cs.CV2024

EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

Xiangpeng Yang, Linchao Zhu, Hehe Fan +1

Current diffusion-based video editing primarily focuses on local editing (\textit{e.g.,} object/background editing) or global style editing by utilizing various dense correspondenc…

cs.CV2026

TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation

Mingwei Li, Hehe Fan, Yi Yang

Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical…

cs.CV2017

Unsupervised Person Re-identification: Clustering and Fine-tuning

Hehe Fan, Liang Zheng, Yi Yang

The superiority of deeply learned pedestrian representations has been reported in very recent literature of person re-identification (re-ID). In this paper, we consider the more pr…

cs.CV2025

InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation

Wenjie Zhuo, Fan Ma, Hehe Fan

We present InfiniDreamer, a novel framework for arbitrarily long human motion generation. InfiniDreamer addresses the limitations of current motion generation methods, which are ty…

cs.RO2026

Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments

Haoyuan Li, Rui Liu, Hehe Fan +1

Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to learn complex reasoning from long-horizon human interactions. While Multi-modal Large Language Mod…

cs.CV2026

Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues

Wenjin Hou, Xiaoxiao Sun, Hehe Fan

Recent advances in zero-shot learning (ZSL) have demonstrated the potential of generative models. Typically, generative ZSL synthesizes visual features conditioned on semantic prot…

cs.CV2025

Learning Robust Representations via Bidirectional Transition for Visual Reinforcement Learning

Xiaobo Hu, Youfang Lin, Yue Liu +4

Visual reinforcement learning has proven effective in solving control tasks with high-dimensional observations. However, extracting reliable and generalizable representations from…

cs.LG2024

Clustering for Protein Representation Learning

Ruijie Quan, Wenguan Wang, Fan Ma +2

Protein representation learning is a challenging task that aims to capture the structure and function of proteins from their amino acid sequences. Previous methods largely ignored…

cs.CV2023

Text to Point Cloud Localization with Relation-Enhanced Transformer

Guangzhi Wang, Hehe Fan, Mohan Kankanhalli

Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, w…

cs.CV2023

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

Zhiqiang Shen, Xiaoxiao Sheng, Hehe Fan +5

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, a…

cs.CV2025

Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding

Haoyuan Li, Rui Liu, Hehe Fan +1

Enabling agents to understand and interact with complex 3D scenes is a fundamental challenge for embodied artificial intelligence systems. While Multimodal Large Language Models (M…

cs.CV2026

Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly

Hang Du, Sicheng Zhang, Binzhu Xie +16

Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial…

cs.CV2020

Adaptive Exploration for Unsupervised Person Re-Identification

Yuhang Ding, Hehe Fan, Mingliang Xu +1

Due to domain bias, directly deploying a deep person re-identification (re-ID) model trained on one dataset often achieves considerably poor accuracy on another dataset. In this pa…

cs.MM2019

Cascaded Revision Network for Novel Object Captioning

Qianyu Feng, Yu Wu, Hehe Fan +2

Image captioning, a challenging task where the machine automatically describes an image by sentences, has drawn significant attention in recent years. Despite the remarkable improv…

cs.CV2024

Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models

Yue Zhang, Hehe Fan, Yi Yang

To bridge the gap between vision and language modalities, Multimodal Large Language Models (MLLMs) usually learn an adapter that converts visual inputs to understandable tokens for…

cs.CV2026

Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction

Yingzhao Jian, Zihao Lin, Hehe Fan

The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However, their prediction results can be inconsi…

cs.CL2026

One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement

Yixiao Zhou, Dongzhou Cheng, zhiliang wu +3

Large Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquiries and the structured logic r…

cs.CV2025

EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space

Jianrong Zhang, Hehe Fan, Yi Yang

Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diff…

cs.CE2026

OSCAgent: Accelerating the Discovery of Organic Solar Cells with LLM Agents

Zhaolin Hu, Zhiliang Wu, Kun Li +2

Organic solar cells (OSCs) hold great promise for sustainable energy, but discovering high-performance materials is time-consuming and costly. Existing molecular generation methods…

cs.CV2024

HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting

Zhenglin Zhou, Fan Ma, Hehe Fan +2

Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods strug…

cs.CV2023

Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering

Yi Cheng, Hehe Fan, Dongyun Lin +3

The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing…

cs.CV2026

MA-Bench: Towards Fine-grained Micro-Action Understanding

Kun Li, Jihao Gu, Fei Wang +3

With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Micro-Action understanding, a vital role in human emotion analysis, remains unexplored du…

cs.CV2023

FlowZero: Zero-Shot Text-to-Video Synthesis with LLM-Driven Dynamic Scene Syntax

Yu Lu, Linchao Zhu, Hehe Fan +1

Text-to-video (T2V) generation is a rapidly growing research area that aims to translate the scenes, objects, and actions within complex video text into a sequence of coherent visu…

cs.CL2026

Super Research: Answering Highly Complex Questions with Large Language Models through Super Deep and Super Wide Research

Yubo Dong, Nianhao You, Yuxuan Hou +5

While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions-those requiring long-horizon plan…

cs.CV2026

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

Kerui Chen, Jinglu Wang, Jianrong Zhang +3

Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…

physics.ins-det2026

Suppression of photon hits in large liquid scintillator detectors via spatiotemporal deep learning

Junle Li, Zhaoxiang Wu, Guanda Gong +5

Liquid scintillator detectors are widely used in neutrino experiments due to their low energy threshold and high energy resolution. Despite the tiny abundance of C in LS, th…

cs.CV2024

VividDreamer: Invariant Score Distillation For Hyper-Realistic Text-to-3D Generation

Wenjie Zhuo, Fan Ma, Hehe Fan +1

This paper presents Invariant Score Distillation (ISD), a novel method for high-fidelity text-to-3D generation. ISD aims to tackle the over-saturation and over-smoothing problems i…

cs.RO2025

Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World

Yingzhao Jian, Zhongan Wang, Yi Yang +1

Humanoid agents often struggle to handle flexible and diverse interactions in open environments. A common solution is to collect massive datasets to train a highly capable model, b…

cs.AI2026

Large language model agents accelerate inverse design of metal-organic frameworks for gas separation

Zhaolin Hu, Hehe Fan, Wangyihan Guo +4

Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design difficult under simultaneo…

cs.CL2025

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs

Yixiao Zhou, Ziyu Zhao, Dongzhou Cheng +6

Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activat…

cs.CV2025

BVINet: Unlocking Blind Video Inpainting with Zero Annotations

Zhiliang Wu, Kerui Chen, Kun Li +2

Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusi…

cs.CV2022

Can We Solve 3D Vision Tasks Starting from A 2D Vision Transformer?

Yi Wang, Zhiwen Fan, Tianlong Chen +2

Vision Transformers (ViTs) have proven to be effective, in solving 2D image understanding tasks by training over large-scale image datasets; and meanwhile as a somehow separate tra…

cs.CV2026

ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation

Kerui Chen, Jianrong Zhang, Ming Li +2

Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion…

cs.CV2026

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO

Fufangchen Zhao, Songbai Tan, Xuerui Qiu +7

Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried in…

cs.CV2026

Deepfake Detection Generalization with Diffusion Noise

Hongyuan Qi, Wenjin Hou, Hehe Fan +1

Deepfake detectors face growing challenges in generalization as new image synthesis techniques emerge. In particular, deepfakes generated by diffusion models are highly photorealis…

cs.LG2026

Variational Rectification Inference for Learning with Noisy Labels

Haoliang Sun, Qi Wei, Lei Feng +4

Label noise has been broadly observed in real-world datasets. To mitigate the negative impact of overfitting to label noise for deep models, effective strategies (\textit{e.g.}, re…

cs.CV2026

Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning

Haoyuan Li, Zhengdong Hu, Jun Wang +2

This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools and exhibit biased tool prefer…

cs.LG2026

CktGen: Automated Analog Circuit Design with Generative Artificial Intelligence

Yuxuan Hou, Hehe Fan, Jianrong Zhang +6

The automatic synthesis of analog circuits presents significant challenges. Most existing approaches formulate the problem as a single-objective optimization task, overlooking that…

cs.CV2026

GraphTARIF: Linear Graph Transformer with Augmented Rank and Improved Focus

Zhaolin Hu, Kun Li, Hehe Fan +1

Linear attention mechanisms have emerged as efficient alternatives to full self-attention in Graph Transformers, offering linear time complexity. However, existing linear attention…

cs.CV2025

Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition

Jihao Gu, Kun Li, Fei Wang +4

Micro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in…

cs.LG2025

Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling

Hehe Fan, Yi Yang, Mohan Kankanhalli +1

When modeling a given type of data, we consider it to involve two key aspects: 1) identifying relevant elements (e.g., image pixels or textual words) to a central element, as in a…

cs.CV2026

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

Wenjin Hou, Wei Liu, Han Hu +3

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. Howeve…

cs.LG2026

A Time-Reparameterized Cumulative Intensity Extrapolation Sampler for Discrete Flow Matching

Feiyang Fu, Hehe Fan

Discrete flow matching (DFM) provides a principled framework for generative modeling on discrete state spaces via continuous-time Markov chain dynamics. In practice, sampling for D…

cs.CV2026

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Yue Zhang, Yingzhao Jian, Yunqiu Xu +2

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric c…

cs.CE2025

ProtChatGPT: Towards Understanding Proteins with Large Language Models

Chao Wang, Hehe Fan, Ruijie Quan +1

Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models…

cs.CV2025

Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion

Zhenglin Zhou, Fan Ma, Hehe Fan +1

Animatable head avatar generation typically requires extensive data for training. To reduce the data requirements, a natural solution is to leverage existing data-free static avata…

cs.CV2026

UniFace: A Unified Fine-grained Face Understanding and Generation Model

Junzhe Li, Sifan Zhou, Liya Guo +9

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and gen…

cs.CV2025

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

Zhiyu Xu, Weilong Yan, Yufei Shi +5

Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific…

cs.CV2023

DPMix: Mixture of Depth and Point Cloud Video Experts for 4D Action Segmentation

Yue Zhang, Hehe Fan, Yi Yang +1

In this technical report, we present our findings from the research conducted on the Human-Object Interaction 4D (HOI4D) dataset for egocentric action segmentation task. As a relat…

cs.CV2023

Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos

Xiaoxiao Sheng, Zhiqiang Shen, Gang Xiao +3

We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at th…

cs.CV2025

VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing

Xiangpeng Yang, Linchao Zhu, Hehe Fan +1

Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level,…

cs.CV2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

Ming Li, Xiangyu Xu, Hehe Fan +7

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks…

cs.CV2022

SEFormer: Structure Embedding Transformer for 3D Object Detection

Xiaoyu Feng, Heming Du, Yueqi Duan +2

Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a key challenge to 3D object detection on point cloud. Recently, Transfo…

cs.CV2026

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection

Hongyuan Qi, Feifei Shao, Ming Li +2

The rapid evolution of video generation technologies poses a significant challenge to media forensics, as conventional detection methods often fail to generalize beyond their train…

cs.CV2025

ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models

Zhenglin Zhou, Fan Ma, Xiaobo Xia +3

We explore inference-time scaling in text-guided 3D diffusion models to enhance generative quality without additional training. To this end, we introduce ITS3D, a framework that fo…

cs.CL2025

Enhancing Large Language Models through Structured Reasoning

Yubo Dong, Hehe Fan

Recent Large Language Models (LLMs) have significantly advanced natural language processing and automated decision-making. However, these models still encounter difficulties when p…

cs.CV2026

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

Junzhe Huang, Xiaoxiao Sun, Yan Yang +6

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluat…

cs.CL2023

DocMSU: A Comprehensive Benchmark for Document-level Multimodal Sarcasm Understanding

Hang Du, Guoshun Nan, Sicheng Zhang +6

Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks an…

cs.CL2025

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

Zhenglin Zhou, Xiaobo Xia, Fan Ma +3

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle…

cs.CV2026

Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models

Ruisi Zhao, Haoren Zheng, Zongxin Yang +2

Rigged 3D assets are fundamental to 3D deformation and animation. However, existing 3D generation methods face challenges in generating animatable geometry, while rigging technique…

cs.CV2023

Continual Learning with Strong Experience Replay

Tao Zhuo, Zhiyong Cheng, Zan Gao +2

Continual Learning (CL) aims at incrementally learning new tasks without forgetting the knowledge acquired from old ones. Experience Replay (ER) is a simple and effective rehearsal…

cs.CV2024

Prototype Learning for Micro-gesture Classification

Guoliang Chen, Fei Wang, Kun Li +5

In this paper, we briefly introduce the solution developed by our team, HFUT-VUT, for the track of Micro-gesture Classification in the MiGA challenge at IJCAI 2024. The task of mic…

cs.CV2024

Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling

Yuze Hao, Jianrong Zhang, Tao Zhuo +2

Hands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotic…

cs.CV2025

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Yue Zhang, Yingzhao Jian, Hehe Fan +2

Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typi…

cs.CV2024

TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

Wei Li, Hehe Fan, Yongkang Wong +2

Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability…

cs.CV2023

Prior-Free Continual Learning with Unlabeled Data in the Wild

Tao Zhuo, Zhiyong Cheng, Hehe Fan +1

Continual Learning (CL) aims to incrementally update a trained model on new tasks without forgetting the acquired knowledge of old ones. Existing CL methods usually reduce forgetti…

cs.CV2025

DeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models

Jianrong Zhang, Hehe Fan, Yi Yang

Human motions are compositional: complex behaviors can be described as combinations of simpler primitives. However, existing approaches primarily focus on forward modeling, e.g., l…

cs.CV2025

Prompt-Aware Controllable Shadow Removal

Kerui Chen, Zhiliang Wu, Wenjin Hou +3

Shadow removal aims to restore the image content in shadowed regions. While deep learning-based methods have shown promising results, they still face key challenges: 1) uncontrolle…

cs.CV2019

Attract or Distract: Exploit the Margin of Open Set

Qianyu Feng, Guoliang Kang, Hehe Fan +1

Open set domain adaptation aims to diminish the domain shift across domains, with partially shared classes. There exist unknown target samples out of the knowledge of source domain…

cs.CV2023

A Study on Differentiable Logic and LLMs for EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2023

Yi Cheng, Ziwei Xu, Fen Fang +5

In this technical report, we present our findings from a study conducted on the EPIC-KITCHENS-100 Unsupervised Domain Adaptation task for Action Recognition. Our research focuses o…

cs.CV2025

TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors

Mingwei Li, Pu Pang, Hehe Fan +2

Reconstructing transparent surfaces is essential for tasks such as robotic manipulation in labs, yet it poses a significant challenge for 3D reconstruction techniques like 3D Gauss…

cs.CV2024

ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning

Wenjin Hou, Dingjie Fu, Kun Li +3

Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen classes to unseen ones, guided by semantic information. To this end, existing…

cs.CV2025

MMAD: Multi-label Micro-Action Detection in Videos

Kun Li, Pengyu Liu, Dan Guo +4

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, whi…

cs.CV2019

PointRNN: Point Recurrent Neural Network for Moving Point Cloud Processing

Hehe Fan, Yi Yang

In this paper, we introduce a Point Recurrent Neural Network (PointRNN) for moving point cloud processing. At each time step, PointRNN takes point coordinates $\boldsymbol{P} \in \…

cs.CV2022

PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences

Hehe Fan, Xin Yu, Yuhang Ding +2

Point cloud sequences are irregular and unordered in the spatial dimension while exhibiting regularities and order in the temporal dimension. Therefore, existing grid based convolu…

cs.CV2019

Cubic LSTMs for Video Prediction

Hehe Fan, Linchao Zhu, Yi Yang

Predicting future frames in videos has become a promising direction of research for both computer vision and robot learning communities. The core of this problem involves moving ob…

cs.LG2026

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Wenjin Hou, Shangpin Peng, Weinong Wang +13

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model…

cs.CV2026

4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

Xindan Zhang, Weilong Yan, Yufei Shi +5

Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing metho…

cs.CV2025

Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

Kun Li, Dan Guo, Guoliang Chen +5

Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for ap…

cs.CV2026

DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification

Ying Shu, Pujian Zhan, Huiqi Yang +3

Both fine-grained discriminative details and global semantic features can contribute to solving person re-identification challenges, such as occlusion and pose variations. Vision f…