papers

Publications (58)

math.OC2025

Tensor Based Proximal Alternating Minimization Method for A Kind of Inhomogeneous Quartic Optimization Problem

Haibin Chen, Yixuan Chen, Chunyan Wang +1

In this paper, we propose an efficient numerical approach for solving a specific type of quartic inhomogeneous polynomial optimization problem inspired by practical applications. T…

cs.AI2026

Understanding and Enforcing Weight Disentanglement in Task Arithmetic

Shangge Liu, Yuehan Yin, Lei Wang +5

Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weig…

cs.CV2026

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning

Jiacheng Yang, Tongying Xiao, Yunkai Dang +7

Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods aim to train models to facili…

cs.CL2025

TimeBill: Time-Budgeted Inference for Large Language Models

Qi Fan, An Zou, Yehan Ma

Large Language Models (LLMs) are increasingly deployed in time-critical systems, such as robotics, autonomous driving, embodied intelligence, and industrial automation, where gener…

cs.CV2020

Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation

Qi Fan, Lei Ke, Wenjie Pei +2

Partially supervised instance segmentation aims to perform learning on limited mask-annotated categories of data thus eliminating expensive and exhaustive mask annotation. The lear…

cs.DB2020

Boosting Cloud Data Analytics using Multi-Objective Optimization

Fei Song, Khaled Zaouk, Chenghao Lyu +4

Data analytics in the cloud has become an integral part of enterprise businesses. Big data analytics systems, however, still lack the ability to take user performance goals and bud…

cs.CV2025

Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank

Shaofeng Zhang, Xuanqi Chen, Ning Liao +7

The dominance of denoising generative models (e.g., diffusion, flow-matching) in visual synthesis is tempered by their substantial training costs and inefficiencies in representati…

cs.DC2024

A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning

Chenghao Lyu, Qi Fan, Philippe Guyard +1

As Spark becomes a common big data analytics platform, its growing complexity makes automatic tuning of numerous parameters critical for performance. Our work on Spark parameter tu…

cs.CV2026

SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models

Zhengxuan Wei, Yi Dong, Zonghui Li +6

Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, existing LoRA merging techniq…

cs.CV2025

FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing

Yunkai Dang, Donghao Wang, Jiacheng Yang +7

Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain…

cs.CV2024

Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities

Qi Fan, Hongyu Yuan, Haolin Zuo +2

Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full mo…

cs.CV2023

Stable Segment Anything Model

Qi Fan, Xin Tao, Lei Ke +6

The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to…

cs.DB2017

VCExplorer: A Interactive Graph Exploration Framework Based on Hub Vertices with Graph Consolidation

Huiju Wang, Zhengkui Wang, Kian-Lee Tan +3

Graphs have been widely used to model different information networks, such as the Web, biological networks and social networks (e.g. Twitter). Due to the size and complexity of the…

cs.SD2024

Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples

Qi Fan, Yutong Li, Yi Xin +3

The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions f…

astro-ph.IM2026

ASCNet: Research on all-sky camera images classification at the Muztagh-ata site

Siqi Wang, Qi Fan, Wenbo Gu +6

Cloud coverage is one of the crucial elements of site testing in astronomy. All-sky camera (ASC) images are beneficial for our research on cloud coverage. In this paper, we propose…

cs.CV2025

Adapting In-Domain Few-Shot Segmentation to New Domains without Source Domain Retraining

Qi Fan, Kaiqi Liu, Nian Liu +4

Cross-domain few-shot segmentation (CD-FSS) aims to segment objects of novel classes in new domains, which is often challenging due to the diverse characteristics of target domains…

cs.CV2026

Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

Junyuan Ma, Xunzhi Xiang, Wenbin Li +2

Vision foundation models (VFMs) have achieved strong performance across various vision tasks. However, it still remains challenging to apply VFMs for cross-domain few-shot segmenta…

cs.CV2026

PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models

Yuanhao Su, Shaofeng Zhang, Xiaosong Jia +1

The development of 3D Vision-Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of pai…

cs.DB2015

Supporting Window Analytics over Large-scale Dynamic Graphs

Qi Fan, Zhengkui Wang, Chee-Yong Chan +1

In relational DBMS, window functions have been widely used to facilitate data analytics. Surprisingly, while similar concepts have been employed for graph analytics, there has been…

cs.CV2026

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework

Jiacheng Yang, Anqi Chen, Yunkai Dang +5

Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resoluti…

cond-mat.supr-con2025

Organometallic-Inorganic Hybrid MXenes with Tunable Superconductivity

Qi Fan, Tao Bo, Wei Guo +19

Ti-based two-dimensional transition-metal carbides (MXenes) have attracted attention due to their superior properties and are being explored across various applications1,2. Despite…

physics.app-ph2024

Gaseous Scissor-mediated Electrochemical Exfoliation of Halogenated MXenes and its Boosting in Wear-Resisting Tribovoltaic Devices

Qi Fan, Minghua Chen, Longyi Li +9

Two-dimensional transition metal carbides (MXenes), especially their few-layered nanosheets, have triggered burgeoning research attentions owing to their superiorities including ex…

cs.CL2024

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97

The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…

cs.CV2026

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning

Jiahua Chen, Qihong Tang, Weinong Wang +1

Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing…

cs.LG2025

Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing Modalities

Rui Liu, Haolin Zuo, Zheng Lian +2

Missing modalities have recently emerged as a critical research direction in multimodal emotion recognition (MER). Conventional approaches typically address this issue through miss…

cs.CV2025

Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation

Xunzhi Xiang, Yabo Chen, Guiyu Zhang +10

Current autoregressive diffusion models excel at video generation but are generally limited to short temporal durations. Our theoretical analysis indicates that the autoregressive…

cs.CV2025

Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs

Yikun Ji, Hong Yan, Jun Lan +5

The rapid advancement of image generation technologies intensifies the demand for interpretable and robust detection methods. Although existing approaches often attain high accurac…

cs.CV2026

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

Xu Guo, Zhengxuan Wei, Xinghui Li +11

Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…

cs.CV2022

Self-Support Few-Shot Semantic Segmentation

Qi Fan, Wenjie Pei, Yu-Wing Tai +1

Existing few-shot segmentation methods have achieved great progress based on the support-query matching framework. But they still heavily suffer from the limited coverage of intra-…

cs.CV2026

Prompt-Free Universal Region Proposal Network

Qihong Tang, Changhan Liu, Shaofeng Zhang +3

Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by…

cs.CV2024

Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios

Qi Fan, Haolin Zuo, Rui Liu +2

Multimodal emotion recognition (MER) in practical scenarios is significantly challenged by the presence of missing or incomplete data across different modalities. To overcome these…

cs.CV2021

Group Collaborative Learning for Co-Salient Object Detection

Qi Fan, Deng-Ping Fan, Huazhu Fu +3

We present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations a…

cs.CV2023

GCoNet+: A Stronger Group Collaborative Co-Salient Object Detector

Peng Zheng, Huazhu Fu, Deng-Ping Fan +5

In this paper, we present a novel end-to-end group collaborative learning network, termed GCoNet+, which can effectively and efficiently (250 fps) identify co-salient objects in na…

cs.CV2025

Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation

Xunzhi Xiang, Qi Fan

Autoregressive conditional image generation models have emerged as a dominant paradigm in text-to-image synthesis. These methods typically convert images into one-dimensional token…

cs.CV2026

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

Shengbin Guo, Shaokang He, Chaoyue Meng +4

While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive evaluation of physics-based re…

cs.LG2026

CUDABench: Benchmarking LLMs for Text-to-CUDA Generation

Jiace Zhu, Wentao Chen, Qi Fan +6

Recent studies have demonstrated the potential of Large Language Models (LLMs) in generating GPU Kernels. Current benchmarks focus on the translation of high-level languages into C…

cs.CV2026

Geometry-Aware Implicit Memory for Video World Models

Zhengxuan Wei, Xu Guo, Xinghui Li +8

Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window…

cs.CV2020

Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector

Qi Fan, Wei Zhuo, Chi-Keung Tang +1

Conventional methods for object detection typically require a substantial amount of training data and preparing such high-quality training data is very labor-intensive. In this pap…

cs.CV2026

DreamWorld: Unified World Modeling in Video Generation

Boming Tan, Xiangdong Zhang, Ning Liao +5

Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior appr…

cs.SI2017

Real-Time Influence Maximization on Dynamic Social Streams

Yanhao Wang, Qi Fan, Yuchen Li +1

Influence maximization (IM), which selects a set of users (called seeds) to maximize the influence spread over a social network, is a fundamental problem in a wide range of app…

cs.CV2026

VMonarch: Efficient Video Diffusion Transformers with Structured Attention

Cheng Liang, Haoxian Chen, Liang Hou +4

The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal a…

cs.CV2022

Normalization Perturbation: A Simple Domain Generalization Method for Real-World Domain Shifts

Qi Fan, Mattia Segu, Yu-Wing Tai +4

Improving model's generalizability against domain shifts is crucial, especially for safety-critical applications such as autonomous driving. Real-world domain styles can vary subst…

cs.DB2022

Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing

Chenghao Lyu, Qi Fan, Fei Song +8

Big data processing at the production scale presents a highly complex environment for resource optimization (RO), a problem crucial for meeting performance goals and budgetary cons…

cs.CV2023

DARNet: Bridging Domain Gaps in Cross-Domain Few-Shot Segmentation with Dynamic Adaptation

Haoran Fan, Qi Fan, Maurice Pagnucco +1

Few-shot segmentation (FSS) aims to segment novel classes in a query image by using only a small number of supporting images from base classes. However, in cross-domain few-shot se…

cs.CV2023

Selective Feature Adapter for Dense Vision Transformers

Xueqing Deng, Qi Fan, Xiaojie Jin +2

Fine-tuning pre-trained transformer models, e.g., Swin Transformer, are successful in numerous downstream for dense prediction vision tasks. However, one major issue is the cost/st…

cs.CV2026

VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning

Zhe Gao, Shiyu Shen, Taifeng Chai +7

Existing Multimodal Large Language Models (MLLMs) often suffer from hallucinations in long video understanding (LVU), primarily due to the imbalance between textual and visual toke…

cs.CV2026

CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models

Yunkai Dang, Yizhu Jiang, Yifan Jiang +4

Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address…

cs.LG2025

CUDA-LLM: LLMs Can Write Efficient CUDA Kernels

Wentao Chen, Jiace Zhu, Qi Fan +2

Large Language Models (LLMs) have demonstrated strong capabilities in general-purpose code generation. However, generating the code which is deeply hardware-specific, architecture-…

cs.CV2026

VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling

Xunzhi Xiang, Zixuan Duan, Yabo Chen +8

Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually e…

cs.LG2024

ONNXPruner: ONNX-Based General Model Pruning Adapter

Dongdong Ren, Wenbin Li, Tianyu Ding +5

Recent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various…

cs.CV2026

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

Ting-Bing Xu, Jiacheng Sui, Zhe Gao +12

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We intr…

cs.CV2026

Pathwise Test-Time Correction for Autoregressive Long Video Generation

Xunzhi Xiang, Zixuan Duan, Guiyu Zhang +7

Distilled autoregressive diffusion models facilitate real-time short video synthesis but suffer from severe error accumulation during long-sequence generation. While existing Test-…

cs.NE2009

Network Topology and Time Criticality Effects in the Modularised Fleet Mix Problem

James M. Whitacre, Axel Bender, Stephen Baker +3

In this paper, we explore the interplay between network topology and time criticality in a military logistics system. A general goal of this work (and previous work) is to evaluate…

cs.CV2025

Denoising Vision Transformer Autoencoder with Spectral Self-Regularization

Xunzhi Xiang, Xingye Tian, Guiyu Zhang +5

Variational autoencoders (VAEs) typically encode images into a compact latent space, reducing computational cost but introducing an optimization dilemma: a higher-dimensional laten…

cs.CV2022

Few-Shot Video Object Detection

Qi Fan, Chi-Keung Tang, Yu-Wing Tai

We introduce Few-Shot Video Object Detection (FSVOD) with three contributions to real-world visual learning challenge in our highly diverse and dynamic world: 1) a large-scale vide…

cs.CV2023

UniBoost: Unsupervised Unimodal Pre-training for Boosting Zero-shot Vision-Language Tasks

Yanan Sun, Zihan Zhong, Qi Fan +2

Large-scale joint training of multimodal models, e.g., CLIP, have demonstrated great performance in many vision-language tasks. However, image-text pairs for pre-training are restr…

cs.CV2024

Domain-Rectifying Adapter for Cross-Domain Few-Shot Segmentation

Jiapeng Su, Qi Fan, Guangming Lu +2

Few-shot semantic segmentation (FSS) has achieved great success on segmenting objects of novel classes, supported by only a few annotated samples. However, existing FSS methods oft…

cs.CV2025

A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs

Yunkai Dang, Meiyi Zhu, Donghao Wang +7

Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on…