7 papers
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Chen-Yi Lu, Yueh-Shao Chen, Somali Chaterji
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to nega…
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Pengcheng Wang, Zhiquan Wang, Jayoung Lee +5
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large…
Digital Guardians: The Past and The Future of Cyber-Physical Resilience
Saurabh Bagchi, Hyunseung Kim, Tarek Abdelzaher +20
Resilience in cyber-physical systems (CPS) is the fundamental ability to maintain safety and critical functionality despite adverse "perturbations," which includes security attacks…
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
Chen Yi Lu, Md Mehrab Tanjim, Ishita Dasgupta +4
We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Lea…
Learning to Inference Adaptively for Multimodal Large Language Models
Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee +4
Multimodal Large Language Models (MLLMs) have shown impressive capabilities in visual reasoning, yet come with substantial computational cost, limiting their deployment in resource…
Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning
Atul Sharma, Kavindu Herath, Saurabh Bagchi +2
We introduce the Hubs and Spokes Learning (HSL) framework, a novel paradigm for collaborative machine learning that combines the strengths of Federated Learning (FL) and Decentrali…