4 papers · 1 filter
Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
Yabin Zhang, Maya Varma, Yunhe Gao +4
Out-of-distribution (OOD) detection aims to identify samples that deviate from in-distribution (ID). One popular pipeline addresses this by introducing negative labels distant from…
A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition
Hongbing Li, Jiamin Liu, Shuo Zhang +1
Grounded Multimodal Named Entity Recognition (GMNER) identifies named entities, including their spans and types, in natural language text and grounds them to the corresponding regi…
DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
Zhen Fang, Zhuoyang Liu, Jiaming Liu +7
To build a generalizable Vision-Language-Action (VLA) model with strong reasoning ability, a common strategy is to first train a specialist VLA on robot demonstrations to acquire r…
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
Chenhongyi Yang, Zehui Chen, Miguel Espinosa +4
We present PlainMamba: a simple non-hierarchical state space model (SSM) designed for general visual recognition. The recent Mamba model has shown how SSMs can be highly competitiv…