42 citations · 213 across the 54 of their papers we have counts for
27 papers · 1 filter
MODAL: Multi-Modal Object Re-ID via Model-Driven Sparse Decoupling and Text-Image Differential Filtering
Chengbo Huang, Jun-Jie Huang, Long Lan +5
Multi-modal object re-identification (Re-ID) aims to facilitate cross-camera object retrieval in complex environments by leveraging complementary information from visual (e.g., RGB…
JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
Ce Chen, Congrui Wang, Yonglin Li +23
Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant challenges for edge deployment in terms of l…
Learning Disentangled Representations for Generalized Multi-view Clustering
Xin Zou, Ruimeng Liu, Chang Tang +4
Multi-View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existing deep MVC methods often st…
Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning
Xin Ning, Qiankun Li, Xiaolong Huang +5
With the accumulation of resources in the era of big data and the rise of pre-trained models in deep learning, optimizing neural networks for various tasks often involves different…
ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language Model
Xiaoshu Chen, Sihang Zhou, Ke Liang +2
Compressing long chains of thought (CoT) into compact latent tokens is crucial for efficient reasoning with large language models (LLMs). Recent studies employ autoencoders to achi…
Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
Xiao He, Chang Tang, Xinwang Liu +5
Hyperspectral images with high spectral resolution provide new insights into recognizing subtle differences in similar substances. However, object detection in hyperspectral images…