105 citations · 146 across the 10 of their papers we have counts for
5 papers · 1 filter
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Chen-Yi Lu, Yueh-Shao Chen, Somali Chaterji
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to nega…
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
Pengcheng Wang, Zhiquan Wang, Jayoung Lee +5
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large…
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
Chen Yi Lu, Md Mehrab Tanjim, Ishita Dasgupta +4
We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Lea…
JANUS: Benchmarking Commercial and Open-Source Cloud and Edge Platforms for Object and Anomaly Detection Workloads
Karthick Shankar, Pengcheng Wang, Ran Xu +2
With diverse IoT workloads, placing compute and analytics close to where data is collected is becoming increasingly important. We seek to understand what is the performance and the…
ApproxDet: Content and Contention-Aware Approximate Object Detection for Mobiles
Ran Xu, Chen-lin Zhang, Pengcheng Wang +5
Advanced video analytic systems, including scene classification and object detection, have seen widespread success in various domains such as smart cities and autonomous transporta…