4 papers
Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics
Minglei Chen, Weilong Wang, Jiang Duan +1
Parameter-efficient prompt learning has become the de facto standard for adapting Vision-Language Models (VLMs) to downstream tasks. Existing approaches predominantly focus on alig…
DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
Yuhan Hao, Zhengning Li, Lei Sun +7
Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level annotation, and evaluation protoc…
A Geometry-Aware Algorithm to Learn Hierarchical Embeddings in Hyperbolic Space
Zhangyu Wang, Lantian Xu, Zhifeng Kong +3
Hyperbolic embeddings are a class of representation learning methods that offer competitive performances when data can be abstracted as a tree-like graph. However, in practice, lea…
BCTR: Bidirectional Conditioning Transformer for Scene Graph Generation
Peng Hao, Weilong Wang, Xiaobing Wang +5
Scene Graph Generation (SGG) remains a challenging task due to its compositional property. Previous approaches improve prediction efficiency through end-to-end learning. However, t…