papers

Publications (14)

cs.CV2026

DrawMotion: Generating 3D Human Motions by Freehand Drawing

Tao Wang, Lei Jin, Zhihua Wu +7

Text-to-motion generation, which translates textual descriptions into human motions, faces the challenge that users often struggle to precisely convey their intended motions throug…

cs.CV2026

Distilling Latent Manifolds: Resolution Extrapolation by Variational Autoencoders

Jiaming Chu, Tao Wang, Lei Jin

Variational Autoencoder (VAE) encoders play a critical role in modern generative models, yet their computational cost often motivates the use of knowledge distillation or quantific…

cs.CV2023

The 3rd Anti-UAV Workshop & Challenge: Methods and Results

Jian Zhao, Jianan Li, Lei Jin +20

The 3rd Anti-UAV Workshop & Challenge aims to encourage research in developing novel and accurate methods for multi-scale object tracking. The Anti-UAV dataset used for the Anti-UA…

cs.CV2025

StickMotion: Generating 3D Human Motions by Drawing a Stickman

Tao Wang, Zhihua Wu, Qiaozhi He +6

Text-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text in…

cs.AI2024

Towards Automated Data Sciences with Natural Language and SageCopilot: Practices and Lessons Learned

Yuan Liao, Jiang Bian, Yuhui Yun +9

While the field of NL2SQL has made significant advancements in translating natural language instructions into executable SQL scripts for data querying and processing, achieving ful…

cs.CV2022

SoccerNet 2022 Challenges Results

Silvio Giancola, Anthony Cioppa, Adrien Deliège +91

The SoccerNet 2022 challenges were the second annual video understanding challenges organized by the SoccerNet team. In 2022, the challenges were composed of 6 vision-based tasks:…

cs.CV2024

UniParser: Multi-Human Parsing with Unified Correlation Representation Learning

Jiaming Chu, Lei Jin, Junliang Xing +1

Multi-human parsing is an image segmentation task necessitating both instance-level and fine-grained category-level information. However, prior research has typically processed the…

cs.CV2026

Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection

Yuchu Jiang, Jiaming Chu, Jian Zhao +5

The proliferation of generative models has raised serious concerns about visual content forgery. Existing deepfake detection methods primarily target either image-level classificat…

cs.CV2025

DiffBrush:Just Painting the Art by Your Hands

Jiaming Chu, Lei Jin, Tao Wang +2

The rapid development of image generation and editing algorithms in recent years has enabled ordinary user to produce realistic images. However, the current AI painting ecosystem p…

cs.CV2023

Single-stage Multi-human Parsing via Point Sets and Center-based Offsets

Jiaming Chu, Lei Jin, Junliang Xing +1

This work studies the multi-human parsing problem. Existing methods, either following top-down or bottom-up two-stage paradigms, usually involve expensive computational costs. We i…

cs.CV2024

Technical Report for SoccerNet Challenge 2022 -- Replay Grounding Task

Shimin Chen, Wei Li, Jiaming Chu +3

In order to make full use of video information, we transform the replay grounding problem into a video action location problem. We apply a unified network Faster-TAD proposed by us…

cs.AI2025

ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection

Xin Zhang, Jiaming Chu, Jian Zhao +5

Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including au…

cs.CV2025

EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

Daiheng Gao, Shilin Lu, Shaw Walters +8

Removing unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is es…

cs.CV2024

MM-SEAL: A Large-scale Video Dataset of Multi-person Multi-grained Spatio-temporally Action Localization

Shimin Chen, Wei Li, Chen Chen +4

In this paper, we introduce a novel large-scale video dataset dubbed MM-SEAL for multi-person multi-grained spatio-temporal action localization among human daily life. We are the f…