9 papers
OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction
Junjie Liu, Wanshui Gan, Zitong Dai +6
3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We…
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
Wenyun Li, Zheng Zhang, Dongmei Jiang +2
Vision Transformer (ViT) models have achieved remarkable performance across various vision tasks, with scalability being a key advantage when applied to large datasets. This scalab…
AlignMamba-2: Enhancing Multimodal Fusion and Sentiment Analysis with Modality-Aware Mamba
Yan Li, Yifei Xing, Xiangyuan Lan +3
In the era of large-scale pre-trained models, effectively adapting general knowledge to specific affective computing tasks remains a challenge, particularly regarding computational…
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
Wenyun Li, Zheng Zhang, Dongmei Jiang +1
Large language models (LLMs) have garnered significant interest in AI community. Despite their impressive generation capabilities, they have been found to produce misleading or fab…
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
Guiping Cao, Xiangyuan Lan, Wenjian Huang +3
Popular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g.,…
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
Guiping Cao, Wenjian Huang, Xiangyuan Lan +3
Small Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promis…