papers

Publications (41)

cs.CV2025

Ref-GS: Directional Factorization for 2D Gaussian Splatting

Youjia Zhang, Anpei Chen, Yumin Wan +4

In this paper, we introduce Ref-GS, a novel approach for directional light factorization in 2D Gaussian splatting, which enables photorealistic view-dependent appearance rendering…

cs.CV2026

GateMOT: Q-Gated Attention for Dense Object Tracking

Mingjin Lv, Zelin Liu, Feifei Shao +4

While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tracking: its quadratic all-to…

cs.CV2023

Dynamic Feature Pruning and Consolidation for Occluded Person Re-Identification

YuTeng Ye, Hang Zhou, Jiale Cai +6

Occluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as huma…

cs.CV2020

Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation

Yawei Luo, Ping Liu, Tao Guan +2

We aim at the problem named One-Shot Unsupervised Domain Adaptation. Unlike traditional Unsupervised Domain Adaptation, it assumes that only one unlabeled target sample can be avai…

cs.CV2019

Significance-aware Information Bottleneck for Domain Adaptive Semantic Segmentation

Yawei Luo, Ping Liu, Tao Guan +2

For unsupervised domain adaptation problems, the strategy of aligning the two domains in latent feature space through adversarial learning has achieved much progress in image class…

cs.CV2024

Autogenic Language Embedding for Coherent Point Tracking

Zikai Song, Ying Tang, Run Luo +4

Point tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences. Recent advancements have primarily focused on te…

cs.CV2023

Compact Transformer Tracker with Correlative Masked Modeling

Zikai Song, Run Luo, Junqing Yu +2

Transformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with t…

cs.CV2025

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding

Yangliu Hu, Zikai Song, Na Feng +4

Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have…

cs.MM2026

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction

Dali Wang, Yunyao Zhang, Junqing Yu +3

Micro-video popularity prediction (MVPP) aims to forecast the future popularity of videos on online media, which is essential for applications such as content recommendation and tr…

cs.AI2026

Semantic-Aware Logical Reasoning via a Semiotic Framework

Yunyao Zhang, Xinglang Zhang, Junxi Sheng +5

Logical reasoning is a fundamental capability of large language models. However, existing studies often overlook the interaction between logical complexity and semantic complexity,…

cs.CV2019

Taking A Closer Look at Domain Shift: Category-level Adversaries for Semantics Consistent Domain Adaptation

Yawei Luo, Liang Zheng, Tao Guan +2

We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distrib…

cs.LG2026

LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing

Wenbing Li, Zikai Song, Hang Zhou +3

Recent attempts to combine low-rank adaptation (LoRA) with mixture-of-experts (MoE) for multi-task adaptation of Large Language Models (LLMs) often replace whole attention/FFN laye…

cs.CV2025

MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation

Qilong Xing, Zikai Song, Youjia Zhang +3

Despite significant advancements in adapting Large Language Models (LLMs) for radiology report generation (RRG), clinical adoption remains challenging due to difficulties in accura…

cs.CV2026

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding

Guiyi Zeng, Junqing Yu, Yi-Ping Phoebe Chen +3

Recent advances in self-evolution video understanding frameworks have demonstrated the potential of autonomous learning without human annotations. However, existing methods often s…

cs.CR2024

Attacking Transformers with Feature Diversity Adversarial Perturbation

Chenxing Gao, Hang Zhou, Junqing Yu +4

Understanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturba tions, is crucial for addressing challenges in its real-world a…

cs.CV2024

TIGER: Text-Instructed 3D Gaussian Retrieval and Coherent Editing

Teng Xu, Jiamin Chen, Peng Chen +3

Editing objects within a scene is a critical functionality required across a broad spectrum of applications in computer vision and graphics. As 3D Gaussian Splatting (3DGS) emerges…

cs.CV2023

AMD:Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion

Beibei Jing, Youjia Zhang, Zikai Song +2

Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both natural language and human motion.R…

cs.CV2022

NeReF: Neural Refractive Field for Fluid Surface Reconstruction and Implicit Representation

Ziyu Wang, Wei Yang, Junming Cao +3

Existing neural reconstruction schemes such as Neural Radiance Field (NeRF) are largely focused on modeling opaque objects. We present a novel neural refractive field(NeReF) to rec…

cs.AI2026

Semiotic logical hexagon theory for LLM logical reasoning

Yunyao Zhang, Xinglang Zhang, Zeliang Chen +2

Large language models (LLMs) have become powerful tools for language understanding and logical reasoning. However, they still make mistakes when a problem requires both understandi…

cs.CV2023

Fine-grained Appearance Transfer with Diffusion Models

Yuteng Ye, Guanwen Li, Hang Zhou +7

Image-to-image translation (I2I), and particularly its subfield of appearance transfer, which seeks to alter the visual appearance between images while maintaining structural coher…

cs.CV2018

Macro-Micro Adversarial Network for Human Parsing

Yawei Luo, Zhedong Zheng, Liang Zheng +3

In human parsing, the pixel-wise classification loss has drawbacks in its low-level local inconsistency and high-level semantic inconsistency. The introduction of the adversarial n…

cs.AI2024

Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model

Wenbing Li, Hang Zhou, Junqing Yu +2

The essence of multi-modal fusion lies in exploiting the complementary information inherent in diverse modalities. However, prevalent fusion methods rely on traditional neural arch…

cs.CV2023

Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection

Hang Zhou, Junqing Yu, Wei Yang

Learning discriminative features for effectively separating abnormal events from normality is crucial for weakly supervised video anomaly detection (WS-VAD) tasks. Existing approac…

cs.CV2025

Optimized View and Geometry Distillation from Multi-view Diffuser

Youjia Zhang, Zikai Song, Junqing Yu +2

Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as…

cs.CV2026

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization

Ying Tang, Dong Li, Youjia Zhang +3

Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent i…

cs.CV2022

Transformer Tracking with Cyclic Shifting Window Attention

Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen +1

Transformer architecture has been showing its great strength in visual object tracking, for its effective attention mechanism. Existing transformer-based approaches adopt the pixel…

cs.CV2025

Cross-Modality Masked Learning for Survival Prediction in ICI Treated NSCLC Patients

Qilong Xing, Zikai Song, Bingxin Gong +3

Accurate prognosis of non-small cell lung cancer (NSCLC) patients undergoing immunotherapy is essential for personalized treatment planning, enabling informed patient decisions, an…

cs.CV2025

MVP: Winning Solution to SMP Challenge 2025 Video Track

Liliang Ye, Yunyao Zhang, Yafeng Wu +4

Social media platforms serve as central hubs for content dissemination, opinion expression, and public engagement across diverse modalities. Accurately predicting the popularity of…

cs.MM2025

HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction

Liliang Ye, Yunyao Zhang, Yafeng Wu +4

Social media popularity prediction plays a crucial role in content optimization, marketing strategies, and user engagement enhancement across digital platforms. However, predicting…

cs.AI2026

HotComment: A Benchmark for Evaluating Popularity of Online Comments

Yafeng Wu, Yunyao Zhang, Liliang Ye +4

Online comments play a crucial role in shaping public sentiment and opinion dynamics on social media. However, evaluating their popularity remains challenging, not only because it…

cs.LG2018

Every Node Counts: Self-Ensembling Graph Convolutional Networks for Semi-Supervised Learning

Yawei Luo, Tao Guan, Junqing Yu +2

Graph convolutional network (GCN) provides a powerful means for graph-based semi-supervised tasks. However, as a localized first-order approximation of spectral graph convolution,…

cs.CV2026

OmniTrend: Content-Context Modeling for Scalable Social Popularity Prediction

Liliang Ye, Guiyi Zeng, Yunyao Zhang +3

Predicting social media popularity requires understanding both the intrinsic appeal of content and the external context that determines how it is exposed to users. Existing methods…

cs.AI2026

Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning

Xinglang Zhang, Yunyao Zhang, ZeLiang Chen +3

Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decision-making in high-stakes domains such…

eess.IV2025

CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation

Qilong Xing, Zikai Song, Yuteng Ye +5

Segmentation of brain structures from MRI is crucial for evaluating brain morphology, yet existing CNN and transformer-based methods struggle to delineate complex structures accura…

cs.SI2026

GA-S: Comprehensive Social Network Simulation with Group Agents

Yunyao Zhang, Zikai Song, Hang Zhou +4

Social network simulation is developed to provide a comprehensive understanding of social networks in the real world, which can be leveraged for a wide range of applications such a…

cs.CV2024

Progressive Text-to-Image Diffusion with Soft Latent Direction

YuTeng Ye, Jiale Cai, Hang Zhou +6

In spite of the rapidly evolving landscape of text-to-image generation, the synthesis and manipulation of multiple entities while adhering to specific relational constraints pose e…

cs.CV2023

NeMF: Inverse Volume Rendering with Neural Microflake Field

Youjia Zhang, Teng Xu, Junqing Yu +5

Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering. Rece…

cs.CV2026

Hypergraph-State Collaborative Reasoning for Multi-Object Tracking

Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen +2

Motion reasoning serves as the cornerstone of multi-object tracking (MOT), as it enables consistent association of targets across frames. However, existing motion estimation approa…

cs.SI2026

IntervenSim: Intervention-Aware Social Network Simulation for Opinion Dynamics

Yunyao Zhang, Zuocheng Ying, Xinglang Zhang +5

LLM-based social network simulation introduces a new computational approach for modeling event evolution in complex online environments. However, existing methods typically simulat…

cs.CV2024

Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model

Hang Zhou, Jiale Cai, Yuteng Ye +5

A recent endeavor in one class of video anomaly detection is to leverage diffusion models and posit the task as a generation problem, where the diffusion model is trained to recove…

cs.SI2026

Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation

Yunyao Zhang, Yihao Ai, Zuocheng Ying +4

Social network simulation aims to model collective opinion dynamics in large populations, but existing LLM-based simulators mainly focus on aggregate dynamics while largely ignorin…