6 papers
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
Ye Li, Huanan Liu, Kangye Ji +7
Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time ro…
Beyond Single Prompts: Synergistic Fusion and Arrangement for VICL
Wenwen Liao, Jianbo Yu, Yuansong Wang +2
Vision In-Context Learning (VICL) enables inpainting models to quickly adapt to new visual tasks from only a few prompts. However, existing methods suffer from two key issues: (1)…
Enhancing Visual In-Context Learning by Multi-Faceted Fusion
Wenwen Liao, Jianbo Yu, Yuansong Wang +2
Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-the…
InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery
Wenwen Liao, Hang Ruan, Jianbo Yu +3
Generalized Category Discovery (GCD) aims to classify instances from both known and novel categories within a large-scale unlabeled dataset, a critical yet challenging task for rea…
EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers
Wenwen Liao, Hang Ruan, Jianbo Yu +3
Large models such as Vision Transformers (ViTs) have demonstrated remarkable superiority over smaller architectures like ResNet in few-shot classification, owing to their powerful…
CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly Detection
Yaohua Zha, Xue Yuerong, Chunlin Fan +4
Deep learning-based 3D anomaly detection methods have demonstrated significant potential in industrial manufacturing. However, many approaches are specifically designed for anomaly…