Publications (83)
AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
Zhenglin Zhou, Fan Ma, Chengzhuo Gui +4
Training-free 3D editing aims to modify 3D shapes based on human instructions without model finetuning. It plays a crucial role in 3D content creation. However, existing approaches…
Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL
Yixiao Zhou, Yang Li, Dongzhou Cheng +2
Reinforcement Learning from Verifiable Rewards (RLVR) trains large language models (LLMs) from sampled trajectories, making decoding strategy a core component of learning rather th…
TV-Dialogue: Crafting Theme-Aware Video Dialogues with Immersive Interaction
Sai Wang, Fan Ma, Xinyi Li +2
Recent advancements in LLMs have accelerated the development of dialogue generation across text and images, yet video-based dialogue generation remains underexplored and presents u…
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
Xiangpeng Yang, Linchao Zhu, Hehe Fan +1
Current diffusion-based video editing primarily focuses on local editing (\textit{e.g.,} object/background editing) or global style editing by utilizing various dense correspondenc…
TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation
Mingwei Li, Hehe Fan, Yi Yang
Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical…
Unsupervised Person Re-identification: Clustering and Fine-tuning
Hehe Fan, Liang Zheng, Yi Yang
The superiority of deeply learned pedestrian representations has been reported in very recent literature of person re-identification (re-ID). In this paper, we consider the more pr…
InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation
Wenjie Zhuo, Fan Ma, Hehe Fan
We present InfiniDreamer, a novel framework for arbitrarily long human motion generation. InfiniDreamer addresses the limitations of current motion generation methods, which are ty…
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
Haoyuan Li, Rui Liu, Hehe Fan +1
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to learn complex reasoning from long-horizon human interactions. While Multi-modal Large Language Mod…
Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
Wenjin Hou, Xiaoxiao Sun, Hehe Fan
Recent advances in zero-shot learning (ZSL) have demonstrated the potential of generative models. Typically, generative ZSL synthesizes visual features conditioned on semantic prot…
Learning Robust Representations via Bidirectional Transition for Visual Reinforcement Learning
Xiaobo Hu, Youfang Lin, Yue Liu +4
Visual reinforcement learning has proven effective in solving control tasks with high-dimensional observations. However, extracting reliable and generalizable representations from…
Clustering for Protein Representation Learning
Ruijie Quan, Wenguan Wang, Fan Ma +2
Protein representation learning is a challenging task that aims to capture the structure and function of proteins from their amino acid sequences. Previous methods largely ignored…
Text to Point Cloud Localization with Relation-Enhanced Transformer
Guangzhi Wang, Hehe Fan, Mohan Kankanhalli
Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, w…
Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos
Zhiqiang Shen, Xiaoxiao Sheng, Hehe Fan +5
Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, a…
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
Haoyuan Li, Rui Liu, Hehe Fan +1
Enabling agents to understand and interact with complex 3D scenes is a fundamental challenge for embodied artificial intelligence systems. While Multimodal Large Language Models (M…
Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
Hang Du, Sicheng Zhang, Binzhu Xie +16
Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial…
Adaptive Exploration for Unsupervised Person Re-Identification
Yuhang Ding, Hehe Fan, Mingliang Xu +1
Due to domain bias, directly deploying a deep person re-identification (re-ID) model trained on one dataset often achieves considerably poor accuracy on another dataset. In this pa…
Cascaded Revision Network for Novel Object Captioning
Qianyu Feng, Yu Wu, Hehe Fan +2
Image captioning, a challenging task where the machine automatically describes an image by sentences, has drawn significant attention in recent years. Despite the remarkable improv…
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
Yue Zhang, Hehe Fan, Yi Yang
To bridge the gap between vision and language modalities, Multimodal Large Language Models (MLLMs) usually learn an adapter that converts visual inputs to understandable tokens for…
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
Yingzhao Jian, Zihao Lin, Hehe Fan
The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However, their prediction results can be inconsi…
One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement
Yixiao Zhou, Dongzhou Cheng, zhiliang wu +3
Large Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquiries and the structured logic r…
EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
Jianrong Zhang, Hehe Fan, Yi Yang
Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diff…
OSCAgent: Accelerating the Discovery of Organic Solar Cells with LLM Agents
Zhaolin Hu, Zhiliang Wu, Kun Li +2
Organic solar cells (OSCs) hold great promise for sustainable energy, but discovering high-performance materials is time-consuming and costly. Existing molecular generation methods…
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
Zhenglin Zhou, Fan Ma, Hehe Fan +2
Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods strug…
Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering
Yi Cheng, Hehe Fan, Dongyun Lin +3
The main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing…
MA-Bench: Towards Fine-grained Micro-Action Understanding
Kun Li, Jihao Gu, Fei Wang +3
With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Micro-Action understanding, a vital role in human emotion analysis, remains unexplored du…
FlowZero: Zero-Shot Text-to-Video Synthesis with LLM-Driven Dynamic Scene Syntax
Yu Lu, Linchao Zhu, Hehe Fan +1
Text-to-video (T2V) generation is a rapidly growing research area that aims to translate the scenes, objects, and actions within complex video text into a sequence of coherent visu…
Super Research: Answering Highly Complex Questions with Large Language Models through Super Deep and Super Wide Research
Yubo Dong, Nianhao You, Yuxuan Hou +5
While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions-those requiring long-horizon plan…
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…
Suppression of photon hits in large liquid scintillator detectors via spatiotemporal deep learning
Junle Li, Zhaoxiang Wu, Guanda Gong +5
Liquid scintillator detectors are widely used in neutrino experiments due to their low energy threshold and high energy resolution. Despite the tiny abundance of C in LS, th…
VividDreamer: Invariant Score Distillation For Hyper-Realistic Text-to-3D Generation
Wenjie Zhuo, Fan Ma, Hehe Fan +1
This paper presents Invariant Score Distillation (ISD), a novel method for high-fidelity text-to-3D generation. ISD aims to tackle the over-saturation and over-smoothing problems i…
Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
Yingzhao Jian, Zhongan Wang, Yi Yang +1
Humanoid agents often struggle to handle flexible and diverse interactions in open environments. A common solution is to collect massive datasets to train a highly capable model, b…
Large language model agents accelerate inverse design of metal-organic frameworks for gas separation
Zhaolin Hu, Hehe Fan, Wangyihan Guo +4
Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes inverse design difficult under simultaneo…
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
Yixiao Zhou, Ziyu Zhao, Dongzhou Cheng +6
Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activat…
BVINet: Unlocking Blind Video Inpainting with Zero Annotations
Zhiliang Wu, Kerui Chen, Kun Li +2
Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusi…
Can We Solve 3D Vision Tasks Starting from A 2D Vision Transformer?
Yi Wang, Zhiwen Fan, Tianlong Chen +2
Vision Transformers (ViTs) have proven to be effective, in solving 2D image understanding tasks by training over large-scale image datasets; and meanwhile as a somehow separate tra…
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
Kerui Chen, Jianrong Zhang, Ming Li +2
Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion…
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO
Fufangchen Zhao, Songbai Tan, Xuerui Qiu +7
Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried in…
Deepfake Detection Generalization with Diffusion Noise
Hongyuan Qi, Wenjin Hou, Hehe Fan +1
Deepfake detectors face growing challenges in generalization as new image synthesis techniques emerge. In particular, deepfakes generated by diffusion models are highly photorealis…
Variational Rectification Inference for Learning with Noisy Labels
Haoliang Sun, Qi Wei, Lei Feng +4
Label noise has been broadly observed in real-world datasets. To mitigate the negative impact of overfitting to label noise for deep models, effective strategies (\textit{e.g.}, re…
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning
Haoyuan Li, Zhengdong Hu, Jun Wang +2
This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools and exhibit biased tool prefer…
CktGen: Automated Analog Circuit Design with Generative Artificial Intelligence
Yuxuan Hou, Hehe Fan, Jianrong Zhang +6
The automatic synthesis of analog circuits presents significant challenges. Most existing approaches formulate the problem as a single-objective optimization task, overlooking that…
GraphTARIF: Linear Graph Transformer with Augmented Rank and Improved Focus
Zhaolin Hu, Kun Li, Hehe Fan +1
Linear attention mechanisms have emerged as efficient alternatives to full self-attention in Graph Transformers, offering linear time complexity. However, existing linear attention…
Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
Jihao Gu, Kun Li, Fei Wang +4
Micro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in…
Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
Hehe Fan, Yi Yang, Mohan Kankanhalli +1
When modeling a given type of data, we consider it to involve two key aspects: 1) identifying relevant elements (e.g., image pixels or textual words) to a central element, as in a…
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
Wenjin Hou, Wei Liu, Han Hu +3
Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. Howeve…
A Time-Reparameterized Cumulative Intensity Extrapolation Sampler for Discrete Flow Matching
Feiyang Fu, Hehe Fan
Discrete flow matching (DFM) provides a principled framework for generative modeling on discrete state spaces via continuous-time Markov chain dynamics. In practice, sampling for D…
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Yue Zhang, Yingzhao Jian, Yunqiu Xu +2
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric c…
ProtChatGPT: Towards Understanding Proteins with Large Language Models
Chao Wang, Hehe Fan, Ruijie Quan +1
Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models…
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
Zhenglin Zhou, Fan Ma, Hehe Fan +1
Animatable head avatar generation typically requires extensive data for training. To reduce the data requirements, a natural solution is to leverage existing data-free static avata…
UniFace: A Unified Fine-grained Face Understanding and Generation Model
Junzhe Li, Sifan Zhou, Liya Guo +9
Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and gen…
SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
Zhiyu Xu, Weilong Yan, Yufei Shi +5
Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific…
DPMix: Mixture of Depth and Point Cloud Video Experts for 4D Action Segmentation
Yue Zhang, Hehe Fan, Yi Yang +1
In this technical report, we present our findings from the research conducted on the Human-Object Interaction 4D (HOI4D) dataset for egocentric action segmentation task. As a relat…
Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos
Xiaoxiao Sheng, Zhiqiang Shen, Gang Xiao +3
We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at th…
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
Xiangpeng Yang, Linchao Zhu, Hehe Fan +1
Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level,…
STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition
Ming Li, Xiangyu Xu, Hehe Fan +7
Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks…
SEFormer: Structure Embedding Transformer for 3D Object Detection
Xiaoyu Feng, Heming Du, Yueqi Duan +2
Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a key challenge to 3D object detection on point cloud. Recently, Transfo…
DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection
Hongyuan Qi, Feifei Shao, Ming Li +2
The rapid evolution of video generation technologies poses a significant challenge to media forensics, as conventional detection methods often fail to generalize beyond their train…
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
Zhenglin Zhou, Fan Ma, Xiaobo Xia +3
We explore inference-time scaling in text-guided 3D diffusion models to enhance generative quality without additional training. To this end, we introduce ITS3D, a framework that fo…
Enhancing Large Language Models through Structured Reasoning
Yubo Dong, Hehe Fan
Recent Large Language Models (LLMs) have significantly advanced natural language processing and automated decision-making. However, these models still encounter difficulties when p…
WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild
Junzhe Huang, Xiaoxiao Sun, Yan Yang +6
Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluat…
DocMSU: A Comprehensive Benchmark for Document-level Multimodal Sarcasm Understanding
Hang Du, Guoshun Nan, Sicheng Zhang +6
Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks an…
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
Zhenglin Zhou, Xiaobo Xia, Fan Ma +3
Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle…
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
Ruisi Zhao, Haoren Zheng, Zongxin Yang +2
Rigged 3D assets are fundamental to 3D deformation and animation. However, existing 3D generation methods face challenges in generating animatable geometry, while rigging technique…
Continual Learning with Strong Experience Replay
Tao Zhuo, Zhiyong Cheng, Zan Gao +2
Continual Learning (CL) aims at incrementally learning new tasks without forgetting the knowledge acquired from old ones. Experience Replay (ER) is a simple and effective rehearsal…
Prototype Learning for Micro-gesture Classification
Guoliang Chen, Fei Wang, Kun Li +5
In this paper, we briefly introduce the solution developed by our team, HFUT-VUT, for the track of Micro-gesture Classification in the MiGA challenge at IJCAI 2024. The task of mic…
Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling
Yuze Hao, Jianrong Zhang, Tao Zhuo +2
Hands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotic…
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
Yue Zhang, Yingzhao Jian, Hehe Fan +2
Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typi…
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
Wei Li, Hehe Fan, Yongkang Wong +2
Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability…
Prior-Free Continual Learning with Unlabeled Data in the Wild
Tao Zhuo, Zhiyong Cheng, Hehe Fan +1
Continual Learning (CL) aims to incrementally update a trained model on new tasks without forgetting the acquired knowledge of old ones. Existing CL methods usually reduce forgetti…
DeMoGen: Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models
Jianrong Zhang, Hehe Fan, Yi Yang
Human motions are compositional: complex behaviors can be described as combinations of simpler primitives. However, existing approaches primarily focus on forward modeling, e.g., l…
Prompt-Aware Controllable Shadow Removal
Kerui Chen, Zhiliang Wu, Wenjin Hou +3
Shadow removal aims to restore the image content in shadowed regions. While deep learning-based methods have shown promising results, they still face key challenges: 1) uncontrolle…
Attract or Distract: Exploit the Margin of Open Set
Qianyu Feng, Guoliang Kang, Hehe Fan +1
Open set domain adaptation aims to diminish the domain shift across domains, with partially shared classes. There exist unknown target samples out of the knowledge of source domain…
A Study on Differentiable Logic and LLMs for EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 2023
Yi Cheng, Ziwei Xu, Fen Fang +5
In this technical report, we present our findings from a study conducted on the EPIC-KITCHENS-100 Unsupervised Domain Adaptation task for Action Recognition. Our research focuses o…
TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors
Mingwei Li, Pu Pang, Hehe Fan +2
Reconstructing transparent surfaces is essential for tasks such as robotic manipulation in labs, yet it poses a significant challenge for 3D reconstruction techniques like 3D Gauss…
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
Wenjin Hou, Dingjie Fu, Kun Li +3
Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen classes to unseen ones, guided by semantic information. To this end, existing…
MMAD: Multi-label Micro-Action Detection in Videos
Kun Li, Pengyu Liu, Dan Guo +4
Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, whi…
PointRNN: Point Recurrent Neural Network for Moving Point Cloud Processing
Hehe Fan, Yi Yang
In this paper, we introduce a Point Recurrent Neural Network (PointRNN) for moving point cloud processing. At each time step, PointRNN takes point coordinates $\boldsymbol{P} \in \…
PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences
Hehe Fan, Xin Yu, Yuhang Ding +2
Point cloud sequences are irregular and unordered in the spatial dimension while exhibiting regularities and order in the temporal dimension. Therefore, existing grid based convolu…
Cubic LSTMs for Video Prediction
Hehe Fan, Linchao Zhu, Yi Yang
Predicting future frames in videos has become a promising direction of research for both computer vision and robot learning communities. The core of this problem involves moving ob…
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Wenjin Hou, Shangpin Peng, Weinong Wang +13
On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model…
4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping
Xindan Zhang, Weilong Yan, Yufei Shi +5
Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing metho…
Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition
Kun Li, Dan Guo, Guoliang Chen +5
Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for ap…
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
Ying Shu, Pujian Zhan, Huiqi Yang +3
Both fine-grained discriminative details and global semantic features can contribute to solving person re-identification challenges, such as occlusion and pose variations. Vision f…