Publications (58)
Tensor Based Proximal Alternating Minimization Method for A Kind of Inhomogeneous Quartic Optimization Problem
Haibin Chen, Yixuan Chen, Chunyan Wang +1
In this paper, we propose an efficient numerical approach for solving a specific type of quartic inhomogeneous polynomial optimization problem inspired by practical applications. T…
Understanding and Enforcing Weight Disentanglement in Task Arithmetic
Shangge Liu, Yuehan Yin, Lei Wang +5
Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weig…
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning
Jiacheng Yang, Tongying Xiao, Yunkai Dang +7
Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods aim to train models to facili…
TimeBill: Time-Budgeted Inference for Large Language Models
Qi Fan, An Zou, Yehan Ma
Large Language Models (LLMs) are increasingly deployed in time-critical systems, such as robotics, autonomous driving, embodied intelligence, and industrial automation, where gener…
Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation
Qi Fan, Lei Ke, Wenjie Pei +2
Partially supervised instance segmentation aims to perform learning on limited mask-annotated categories of data thus eliminating expensive and exhaustive mask annotation. The lear…
Boosting Cloud Data Analytics using Multi-Objective Optimization
Fei Song, Khaled Zaouk, Chenghao Lyu +4
Data analytics in the cloud has become an integral part of enterprise businesses. Big data analytics systems, however, still lack the ability to take user performance goals and bud…
Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
Shaofeng Zhang, Xuanqi Chen, Ning Liao +7
The dominance of denoising generative models (e.g., diffusion, flow-matching) in visual synthesis is tempered by their substantial training costs and inefficiencies in representati…
A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning
Chenghao Lyu, Qi Fan, Philippe Guyard +1
As Spark becomes a common big data analytics platform, its growing complexity makes automatic tuning of numerous parameters critical for performance. Our work on Spark parameter tu…
SSR-Merge: Subspace Signal Routing for Training-Free LoRA Merging in Diffusion Models
Zhengxuan Wei, Yi Dong, Zonghui Li +6
Low-Rank Adaptation (LoRA) merging can efficiently combine diverse generative capabilities from multiple trained LoRAs for a diffusion model. However, existing LoRA merging techniq…
FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing
Yunkai Dang, Donghao Wang, Jiacheng Yang +7
Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain…
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
Qi Fan, Hongyu Yuan, Haolin Zuo +2
Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full mo…
Stable Segment Anything Model
Qi Fan, Xin Tao, Lei Ke +6
The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to…
VCExplorer: A Interactive Graph Exploration Framework Based on Hub Vertices with Graph Consolidation
Huiju Wang, Zhengkui Wang, Kian-Lee Tan +3
Graphs have been widely used to model different information networks, such as the Web, biological networks and social networks (e.g. Twitter). Due to the size and complexity of the…
Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples
Qi Fan, Yutong Li, Yi Xin +3
The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions f…
ASCNet: Research on all-sky camera images classification at the Muztagh-ata site
Siqi Wang, Qi Fan, Wenbo Gu +6
Cloud coverage is one of the crucial elements of site testing in astronomy. All-sky camera (ASC) images are beneficial for our research on cloud coverage. In this paper, we propose…
Adapting In-Domain Few-Shot Segmentation to New Domains without Source Domain Retraining
Qi Fan, Kaiqi Liu, Nian Liu +4
Cross-domain few-shot segmentation (CD-FSS) aims to segment objects of novel classes in new domains, which is often challenging due to the diverse characteristics of target domains…
Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
Junyuan Ma, Xunzhi Xiang, Wenbin Li +2
Vision foundation models (VFMs) have achieved strong performance across various vision tasks. However, it still remains challenging to apply VFMs for cross-domain few-shot segmenta…
PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
Yuanhao Su, Shaofeng Zhang, Xiaosong Jia +1
The development of 3D Vision-Language Models (VLMs), crucial for applications in robotics, autonomous driving, and augmented reality, is severely constrained by the scarcity of pai…
Supporting Window Analytics over Large-scale Dynamic Graphs
Qi Fan, Zhengkui Wang, Chee-Yong Chan +1
In relational DBMS, window functions have been widely used to facilitate data analytics. Surprisingly, while similar concepts have been employed for graph analytics, there has been…
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
Jiacheng Yang, Anqi Chen, Yunkai Dang +5
Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resoluti…
Organometallic-Inorganic Hybrid MXenes with Tunable Superconductivity
Qi Fan, Tao Bo, Wei Guo +19
Ti-based two-dimensional transition-metal carbides (MXenes) have attracted attention due to their superior properties and are being explored across various applications1,2. Despite…
Gaseous Scissor-mediated Electrochemical Exfoliation of Halogenated MXenes and its Boosting in Wear-Resisting Tribovoltaic Devices
Qi Fan, Minghua Chen, Longyi Li +9
Two-dimensional transition metal carbides (MXenes), especially their few-layered nanosheets, have triggered burgeoning research attentions owing to their superiorities including ex…
InternLM2 Technical Report
Zheng Cai, Maosong Cao, Haojiong Chen +97
The evolution of Large Language Models (LLMs) like ChatGPT and GPT-4 has sparked discussions on the advent of Artificial General Intelligence (AGI). However, replicating such advan…
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
Jiahua Chen, Qihong Tang, Weinong Wang +1
Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing…
Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing Modalities
Rui Liu, Haolin Zuo, Zheng Lian +2
Missing modalities have recently emerged as a critical research direction in multimodal emotion recognition (MER). Conventional approaches typically address this issue through miss…
Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
Xunzhi Xiang, Yabo Chen, Guiyu Zhang +10
Current autoregressive diffusion models excel at video generation but are generally limited to short temporal durations. Our theoretical analysis indicates that the autoregressive…
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
Yikun Ji, Hong Yan, Jun Lan +5
The rapid advancement of image generation technologies intensifies the demand for interpretable and robust detection methods. Although existing approaches often attain high accurac…
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Xu Guo, Zhengxuan Wei, Xinghui Li +11
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs.…
Self-Support Few-Shot Semantic Segmentation
Qi Fan, Wenjie Pei, Yu-Wing Tai +1
Existing few-shot segmentation methods have achieved great progress based on the support-query matching framework. But they still heavily suffer from the limited coverage of intra-…
Prompt-Free Universal Region Proposal Network
Qihong Tang, Changhan Liu, Shaofeng Zhang +3
Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by…
Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios
Qi Fan, Haolin Zuo, Rui Liu +2
Multimodal emotion recognition (MER) in practical scenarios is significantly challenged by the presence of missing or incomplete data across different modalities. To overcome these…
Group Collaborative Learning for Co-Salient Object Detection
Qi Fan, Deng-Ping Fan, Huazhu Fu +3
We present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations a…
GCoNet+: A Stronger Group Collaborative Co-Salient Object Detector
Peng Zheng, Huazhu Fu, Deng-Ping Fan +5
In this paper, we present a novel end-to-end group collaborative learning network, termed GCoNet+, which can effectively and efficiently (250 fps) identify co-salient objects in na…
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
Xunzhi Xiang, Qi Fan
Autoregressive conditional image generation models have emerged as a dominant paradigm in text-to-image synthesis. These methods typically convert images into one-dimensional token…
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
Shengbin Guo, Shaokang He, Chaoyue Meng +4
While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive evaluation of physics-based re…
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
Jiace Zhu, Wentao Chen, Qi Fan +6
Recent studies have demonstrated the potential of Large Language Models (LLMs) in generating GPU Kernels. Current benchmarks focus on the translation of high-level languages into C…
Geometry-Aware Implicit Memory for Video World Models
Zhengxuan Wei, Xu Guo, Xinghui Li +8
Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window…
Few-Shot Object Detection with Attention-RPN and Multi-Relation Detector
Qi Fan, Wei Zhuo, Chi-Keung Tang +1
Conventional methods for object detection typically require a substantial amount of training data and preparing such high-quality training data is very labor-intensive. In this pap…
DreamWorld: Unified World Modeling in Video Generation
Boming Tan, Xiangdong Zhang, Ning Liao +5
Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior appr…
Real-Time Influence Maximization on Dynamic Social Streams
Yanhao Wang, Qi Fan, Yuchen Li +1
Influence maximization (IM), which selects a set of users (called seeds) to maximize the influence spread over a social network, is a fundamental problem in a wide range of app…
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
Cheng Liang, Haoxian Chen, Liang Hou +4
The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal a…
Normalization Perturbation: A Simple Domain Generalization Method for Real-World Domain Shifts
Qi Fan, Mattia Segu, Yu-Wing Tai +4
Improving model's generalizability against domain shifts is crucial, especially for safety-critical applications such as autonomous driving. Real-world domain styles can vary subst…
Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing
Chenghao Lyu, Qi Fan, Fei Song +8
Big data processing at the production scale presents a highly complex environment for resource optimization (RO), a problem crucial for meeting performance goals and budgetary cons…
DARNet: Bridging Domain Gaps in Cross-Domain Few-Shot Segmentation with Dynamic Adaptation
Haoran Fan, Qi Fan, Maurice Pagnucco +1
Few-shot segmentation (FSS) aims to segment novel classes in a query image by using only a small number of supporting images from base classes. However, in cross-domain few-shot se…
Selective Feature Adapter for Dense Vision Transformers
Xueqing Deng, Qi Fan, Xiaojie Jin +2
Fine-tuning pre-trained transformer models, e.g., Swin Transformer, are successful in numerous downstream for dense prediction vision tasks. However, one major issue is the cost/st…
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
Zhe Gao, Shiyu Shen, Taifeng Chai +7
Existing Multimodal Large Language Models (MLLMs) often suffer from hallucinations in long video understanding (LVU), primarily due to the imbalance between textual and visual toke…
CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models
Yunkai Dang, Yizhu Jiang, Yifan Jiang +4
Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address…
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
Wentao Chen, Jiace Zhu, Qi Fan +2
Large Language Models (LLMs) have demonstrated strong capabilities in general-purpose code generation. However, generating the code which is deeply hardware-specific, architecture-…
VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling
Xunzhi Xiang, Zixuan Duan, Yabo Chen +8
Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually e…
ONNXPruner: ONNX-Based General Model Pruning Adapter
Dongdong Ren, Wenbin Li, Tianyu Ding +5
Recent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various…
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
Ting-Bing Xu, Jiacheng Sui, Zhe Gao +12
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We intr…
Pathwise Test-Time Correction for Autoregressive Long Video Generation
Xunzhi Xiang, Zixuan Duan, Guiyu Zhang +7
Distilled autoregressive diffusion models facilitate real-time short video synthesis but suffer from severe error accumulation during long-sequence generation. While existing Test-…
Network Topology and Time Criticality Effects in the Modularised Fleet Mix Problem
James M. Whitacre, Axel Bender, Stephen Baker +3
In this paper, we explore the interplay between network topology and time criticality in a military logistics system. A general goal of this work (and previous work) is to evaluate…
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
Xunzhi Xiang, Xingye Tian, Guiyu Zhang +5
Variational autoencoders (VAEs) typically encode images into a compact latent space, reducing computational cost but introducing an optimization dilemma: a higher-dimensional laten…
Few-Shot Video Object Detection
Qi Fan, Chi-Keung Tang, Yu-Wing Tai
We introduce Few-Shot Video Object Detection (FSVOD) with three contributions to real-world visual learning challenge in our highly diverse and dynamic world: 1) a large-scale vide…
UniBoost: Unsupervised Unimodal Pre-training for Boosting Zero-shot Vision-Language Tasks
Yanan Sun, Zihan Zhong, Qi Fan +2
Large-scale joint training of multimodal models, e.g., CLIP, have demonstrated great performance in many vision-language tasks. However, image-text pairs for pre-training are restr…
Domain-Rectifying Adapter for Cross-Domain Few-Shot Segmentation
Jiapeng Su, Qi Fan, Guangming Lu +2
Few-shot semantic segmentation (FSS) has achieved great success on segmenting objects of novel classes, supported by only a few annotated samples. However, existing FSS methods oft…
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
Yunkai Dang, Meiyi Zhu, Donghao Wang +7
Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on…