Publications (41)
Ref-GS: Directional Factorization for 2D Gaussian Splatting
Youjia Zhang, Anpei Chen, Yumin Wan +4
In this paper, we introduce Ref-GS, a novel approach for directional light factorization in 2D Gaussian splatting, which enables photorealistic view-dependent appearance rendering…
GateMOT: Q-Gated Attention for Dense Object Tracking
Mingjin Lv, Zelin Liu, Feifei Shao +4
While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tracking: its quadratic all-to…
Dynamic Feature Pruning and Consolidation for Occluded Person Re-Identification
YuTeng Ye, Hang Zhou, Jiale Cai +6
Occluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as huma…
Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation
Yawei Luo, Ping Liu, Tao Guan +2
We aim at the problem named One-Shot Unsupervised Domain Adaptation. Unlike traditional Unsupervised Domain Adaptation, it assumes that only one unlabeled target sample can be avai…
Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation
Yawei Luo, Ping Liu, Tao Guan +2
For unsupervised domain adaptation problems, the strategy of aligning the two domains in latent feature space through adversarial learning has achieved much progress in image class…
Autogenic Language Embedding for Coherent Point Tracking
Zikai Song, Ying Tang, Run Luo +4
Point tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences. Recent advancements have primarily focused on te…
Compact Transformer Tracker with Correlative Masked Modeling
Zikai Song, Run Luo, Junqing Yu +2
Transformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with t…
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
Yangliu Hu, Zikai Song, Na Feng +4
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have…
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction
Dali Wang, Yunyao Zhang, Junqing Yu +3
Micro-video popularity prediction (MVPP) aims to forecast the future popularity of videos on online media, which is essential for applications such as content recommendation and tr…
Semantic-Aware Logical Reasoning via a Semiotic Framework
Yunyao Zhang, Xinglang Zhang, Junxi Sheng +5
Logical reasoning is a fundamental capability of large language models. However, existing studies often overlook the interaction between logical complexity and semantic complexity,…
Taking A Closer Look at Domain Shift: Category-level Adversaries for Semantics Consistent Domain Adaptation
Yawei Luo, Liang Zheng, Tao Guan +2
We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distrib…
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
Wenbing Li, Zikai Song, Hang Zhou +3
Recent attempts to combine low-rank adaptation (LoRA) with mixture-of-experts (MoE) for multi-task adaptation of Large Language Models (LLMs) often replace whole attention/FFN laye…
MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation
Qilong Xing, Zikai Song, Youjia Zhang +3
Despite significant advancements in adapting Large Language Models (LLMs) for radiology report generation (RRG), clinical adoption remains challenging due to difficulties in accura…
CurEvo: Curriculum-Guided Self-Evolution for Video Understanding
Guiyi Zeng, Junqing Yu, Yi-Ping Phoebe Chen +3
Recent advances in self-evolution video understanding frameworks have demonstrated the potential of autonomous learning without human annotations. However, existing methods often s…
Attacking Transformers with Feature Diversity Adversarial Perturbation
Chenxing Gao, Hang Zhou, Junqing Yu +4
Understanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturba tions, is crucial for addressing challenges in its real-world a…
TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing
Teng Xu, Jiamin Chen, Peng Chen +3
Editing objects within a scene is a critical functionality required across a broad spectrum of applications in computer vision and graphics. As 3D Gaussian Splatting (3DGS) emerges…
AMD:Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion
Beibei Jing, Youjia Zhang, Zikai Song +2
Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both natural language and human motion.R…
NeReF: Neural Refractive Field for Fluid Surface Reconstruction and Implicit Representation
Ziyu Wang, Wei Yang, Junming Cao +3
Existing neural reconstruction schemes such as Neural Radiance Field (NeRF) are largely focused on modeling opaque objects. We present a novel neural refractive field(NeReF) to rec…
Semiotic logical hexagon theory for LLM logical reasoning
Yunyao Zhang, Xinglang Zhang, Zeliang Chen +2
Large language models (LLMs) have become powerful tools for language understanding and logical reasoning. However, they still make mistakes when a problem requires both understandi…
Fine-grained Appearance Transfer with Diffusion Models
Yuteng Ye, Guanwen Li, Hang Zhou +7
Image-to-image translation (I2I), and particularly its subfield of appearance transfer, which seeks to alter the visual appearance between images while maintaining structural coher…
Macro-Micro Adversarial Network for Human Parsing
Yawei Luo, Zhedong Zheng, Liang Zheng +3
In human parsing, the pixel-wise classification loss has drawbacks in its low-level local inconsistency and high-level semantic inconsistency. The introduction of the adversarial n…
Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model
Wenbing Li, Hang Zhou, Junqing Yu +2
The essence of multi-modal fusion lies in exploiting the complementary information inherent in diverse modalities. However, prevalent fusion methods rely on traditional neural arch…
Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection
Hang Zhou, Junqing Yu, Wei Yang
Learning discriminative features for effectively separating abnormal events from normality is crucial for weakly supervised video anomaly detection (WS-VAD) tasks. Existing approac…
Optimized View and Geometry Distillation from Multi-view Diffuser
Youjia Zhang, Zikai Song, Junqing Yu +2
Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as…
PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization
Ying Tang, Dong Li, Youjia Zhang +3
Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent i…
Transformer Tracking with Cyclic Shifting Window Attention
Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen +1
Transformer architecture has been showing its great strength in visual object tracking, for its effective attention mechanism. Existing transformer-based approaches adopt the pixel…
Cross-Modality Masked Learning for Survival Prediction in ICI Treated NSCLC Patients
Qilong Xing, Zikai Song, Bingxin Gong +3
Accurate prognosis of non-small cell lung cancer (NSCLC) patients undergoing immunotherapy is essential for personalized treatment planning, enabling informed patient decisions, an…
MVP: Winning Solution to SMP Challenge 2025 Video Track
Liliang Ye, Yunyao Zhang, Yafeng Wu +4
Social media platforms serve as central hubs for content dissemination, opinion expression, and public engagement across diverse modalities. Accurately predicting the popularity of…
HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction
Liliang Ye, Yunyao Zhang, Yafeng Wu +4
Social media popularity prediction plays a crucial role in content optimization, marketing strategies, and user engagement enhancement across digital platforms. However, predicting…
HotComment: A Benchmark for Evaluating Popularity of Online Comments
Yafeng Wu, Yunyao Zhang, Liliang Ye +4
Online comments play a crucial role in shaping public sentiment and opinion dynamics on social media. However, evaluating their popularity remains challenging, not only because it…
Every Node Counts: Self-Ensembling Graph Convolutional Networks for Semi-Supervised Learning
Yawei Luo, Tao Guan, Junqing Yu +2
Graph convolutional network (GCN) provides a powerful means for graph-based semi-supervised tasks. However, as a localized first-order approximation of spectral graph convolution,…
OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction
Liliang Ye, Guiyi Zeng, Yunyao Zhang +3
Predicting social media popularity requires understanding both the intrinsic appeal of content and the external context that determines how it is exposed to users. Existing methods…
Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning
Xinglang Zhang, Yunyao Zhang, ZeLiang Chen +3
Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decision-making in high-stakes domains such…
CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation
Qilong Xing, Zikai Song, Yuteng Ye +5
Segmentation of brain structures from MRI is crucial for evaluating brain morphology, yet existing CNN and transformer-based methods struggle to delineate complex structures accura…
GA-S: Comprehensive Social Network Simulation with Group Agents
Yunyao Zhang, Zikai Song, Hang Zhou +4
Social network simulation is developed to provide a comprehensive understanding of social networks in the real world, which can be leveraged for a wide range of applications such a…
Progressive Text-to-Image Diffusion with Soft Latent Direction
YuTeng Ye, Jiale Cai, Hang Zhou +6
In spite of the rapidly evolving landscape of text-to-image generation, the synthesis and manipulation of multiple entities while adhering to specific relational constraints pose e…
NeMF: Inverse Volume Rendering with Neural Microflake Field
Youjia Zhang, Teng Xu, Junqing Yu +5
Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering. Rece…
Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen +2
Motion reasoning serves as the cornerstone of multi-object tracking (MOT), as it enables consistent association of targets across frames. However, existing motion estimation approa…
IntervenSim: Intervention-Aware Social Network Simulation for Opinion Dynamics
Yunyao Zhang, Zuocheng Ying, Xinglang Zhang +5
LLM-based social network simulation introduces a new computational approach for modeling event evolution in complex online environments. However, existing methods typically simulat…
Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
Hang Zhou, Jiale Cai, Yuteng Ye +5
A recent endeavor in one class of video anomaly detection is to leverage diffusion models and posit the task as a generation problem, where the diffusion model is trained to recove…
Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation
Yunyao Zhang, Yihao Ai, Zuocheng Ying +4
Social network simulation aims to model collective opinion dynamics in large populations, but existing LLM-based simulators mainly focus on aggregate dynamics while largely ignorin…