6 papers
LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection
Tianfang Zhang, Fengyi Wu, Lei Li +4
Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamental task in remote sensing. Deep learning m…
RPCANet++: Deep Interpretable Robust PCA for Sparse Object Segmentation
Fengyi Wu, Yimian Dai, Tianfang Zhang +4
Robust principal component analysis (RPCA) decomposes an observation matrix into low-rank background and sparse object components. This capability has enabled its application in ta…
Bayesian Optimization for Controlled Image Editing via LLMs
Chengkun Cai, Haoliang Liu, Xu Zhao +6
In the rapidly evolving field of image generation, achieving precise control over generated content and maintaining semantic consistency remain significant limitations, particularl…
The Role of Deductive and Inductive Reasoning in Large Language Models
Chengkun Cai, Xu Zhao, Haoliang Liu +5
Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning tasks, yet their reliance on static prompt structures and limited adaptability to complex scenar…
Human Motion Instruction Tuning
Lei Li, Sen Jia, Jianhao Wang +6
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning ap…
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
Tianfang Zhang, Lei Li, Yang Zhou +4
Vision Transformers (ViTs) mark a revolutionary advance in neural networks with their token mixer's powerful global context capability. However, the pairwise token affinity and com…