Publications (36)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
Yating Liu, Yaowei Li, Zimo Liu +3
Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and lan…
NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images
Lingen Li, Zhaoyang Zhang, Yaowei Li +7
Recent advancements in generative models have significantly improved novel view synthesis (NVS) from multi-view data. However, existing methods depend on external multi-view alignm…
Self Gradient Forcing: Native Long Video Extrapolation
Junhao Zhuang, Shiyi Zhang, Yuxuan Bian +11
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-tru…
The FAST Discovery of a Millisecond Pulsar M15O (PSR J2129+1210O) Hidden in the Harmonics of M15A (PSR J2129+1210A)
Yinfeng Dai, Zhichen Pan, Lei Qian +6
We report the discovery of an isolated millisecond pulsar M15O (J2129+1210O) from the globular cluster M15 (NGC 7078) with a period of 11.06686 ms and a dispersion measure of…
Image Conductor: Precision Control for Interactive Video Synthesis
Yaowei Li, Xintao Wang, Zhaoyang Zhang +5
Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world…
Search for Periodic Radio Signals from Double Neutron Star System Companions Using the Fast Folding Algorithm
Wenze Li, Zhichen Pan, Lei Qian +14
As most of the companions in the double neutron star systems should be normal pulsars, the Fast Folding Algorithm (FFA), which is suitable for finding these long spin period pulsar…
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
Shuzhou Yang, Xiaodong Cun, Xiaoyu Li +2
Given the high complexity of directly generating high-dimensional data such as 4D, we present 4DVD, a cascaded video diffusion model that generates 4D content in a decoupled manner…
IC-Custom: Diverse Image Customization via In-Context Learning
Yaowei Li, Xiaoyu Li, Zhaoyang Zhang +11
Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventiona…
FAST Discovery of Eight Isolated Millisecond Pulsars in NGC 6517
Dejiang Yin, Li-yun Zhang, Lei Qian +14
We present the discovery of 8 isolated millisecond pulsars in Globular Cluster (GC) NGC 6517 using the Five-Hundred-meter Aperture Spherical radio Telescope (FAST). The spin period…
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory
Hongxiang Li, Meng Cao, Xuxin Cheng +3
The recent video grounding works attempt to introduce vanilla contrastive learning into video grounding. However, we claim that this naive solution is suboptimal. Contrastive learn…
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
Yaowei Li, Lingen Li, Zhaoyang Zhang +6
As user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-b…
Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
Yuxuan Bian, Zeyue Xue, Songchun Zhang +9
We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and…
The discovery of three pulsars in the globular cluster M15 with the FAST
Yuxiao Wu, Zhichen Pan, Lei Qian +17
We present the discovery of three pulsars in the Globular Cluster (GC) M15 (NGC 7078) by the Five-hundred-meter Aperture Spherical radio Telescope (FAST). PSR J2129+1210J (M15J) is…
Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation
Yaowei Li, Bang Yang, Xuxin Cheng +3
Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establis…
DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval
Yating Liu, Zimo Liu, Xiangyuan Lan +3
Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person…
Illuminating Hidden Pulsars: Scintillation-Enhanced Discovery of Two Binary Millisecond Pulsars in M13 with FAST
Dejiang Yin, Lin Wang, Li-yun Zhang +7
We conducted a sensitive acceleration search using Fast Fourier Transform (FFT) techniques on full-length and segmented data from 84 observations of the globular cluster M13 with t…
Timing and Scintillation Studies of Pulsars in Globular Cluster M3 (NGC 5272) with FAST
Baoda Li, Li-yun Zhang, Jumei Yao +16
We present the phase-connected timing solutions of all the five pulsars in globular cluster (GC) M3 (NGC 5272), namely PSRs M3A to F (PSRs J1342+2822A to F), with the exception of…
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
Hongxiang Li, Yaowei Li, Yuhang Yang +4
Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleto…
Exploiting Auxiliary Caption for Video Grounding
Hongxiang Li, Meng Cao, Xuxin Cheng +3
Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the {sparsity dilemma} in video annotations, wh…
UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
Yating Liu, Yaowei Li, Xiangyuan Lan +3
Text-based Person Retrieval (TPR) as a multi-modal task, which aims to retrieve the target person from a pool of candidate images given a text description, has recently garnered co…
The Stack Search Tests on FAST Data: Discovery of Six Faint Isolated Millisecond Pulsars in NGC 6517 and NGC 7078 (M15)
Yinfeng Dai, Xing-Jiang Zhu, Zhichen Pan +16
The paper reports the discovery of six faint, isolated millisecond pulsars in the globular clusters NGC 6517 and M15 using the FAST radio telescope, achieved by stacking power spec…
ClimateIQA: A New Dataset and Benchmark to Advance Vision-Language Models in Meteorology Anomalies Analysis
Jian Chen, Peilin Zhou, Yining Hua +7
Meteorological heatmaps play a vital role in deciphering extreme weather phenomena, yet their inherent complexities marked by irregular contours, unstructured patterns, and complex…
The FAST Globular Cluster Pulsar Survey (GC FANS)
Yujie Lian, Zhichen Pan, Haiyan Zhang +24
By January 2025, 60 pulsars were discovered by the Five-hundred-meter Aperture Spherical radio Telescope globular cluster (GC) pulsar survey (GC FANS), with spin periods spanning 1…
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
Hongxiang Li, Yaowei Li, Bin Lin +7
Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal inte…
Efficient Multimodal Fusion via Interactive Prompting
Yaowei Li, Ruijie Quan, Linchao Zhu +1
Large-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multi-modal learning mo…
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
Bang Yang, Yong Dai, Xuxin Cheng +3
While vision-language pre-trained models (VL-PTMs) have advanced multimodal research in recent years, their mastery in a few languages like English restricts their applicability in…
SSVMR: Saliency-based Self-training for Video-Music Retrieval
Xuxin Cheng, Zhihong Zhu, Hongxiang Li +2
With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws…
FAST Observation and Results for Core Collapse Globular Cluster M15 and NGC 6517
Yuxiao Wu, Dejiang Yin, Yu Pan +9
Radio astronomy is part of radio science that developed rapidly in recent decades. In the research of radio astronomy, pulsars have always been an enduring popular research target.…
High Quality Underwater Image Compression with Adaptive Color Correction
Yimin Zhou, Yichong Xia, Sicheng Pan +6
With the increasing exploration and exploitation of the underwater world, underwater images have become a critical medium for human interaction with marine environments, driving ex…
SIAD: Self-supervised Image Anomaly Detection System
Jiawei Li, Chenxi Lan, Xinyi Zhang +7
Recent trends in AIGC effectively boosted the application of visual inspection. However, most of the available systems work in a human-in-the-loop manner and can not provide long-t…
Echo-Memory: A Controlled Study of Memory in Action World Models
Wayne King, Zeyue Xue, Yuxuan Bian +13
We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models. These models generate multi-segment videos from a first frame, text pro…
Searching for pulsars in Globular Clusters with the Fast Fold Algorithm and a new pulsar discovered in M13
Yaowei Li, Lin Wang, Lei Qian +15
We employed the Fast Folding Algorithm (FFA) on L-Band Globular Cluster (GC) observations taken with Five-hundred-meter Aperture Spherical radio Telescope (FAST) to search for new…
Millisecond Pulsars in M2: New discoveries and a detailed timing analysis
Baoda Li, Kuo Liu, Lin Wang +13
Globular clusters (GCs) offer a unique environment for discovering and studying millisecond pulsars. In this paper, we present a multi-epoch search and detailed timing analysis of…
BrushEdit: All-In-One Image Inpainting and Editing
Yaowei Li, Yuxuan Bian, Xuan Ju +5
Image editing has advanced significantly with the development of diffusion models using both inversion-based and instruction-based methods. However, current inversion-based approac…
Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions
Luxury, Jie Huang, Zihao Fan +25
While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-tim…
ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
Lingen Li, Guangzhi Wang, Zhaoyang Zhang +6
Traditional cartoon and anime production involves keyframing, inbetweening, and colorization stages, which require intensive manual effort. Despite recent advances in AI, existing…