15 papers
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
Woo Jae Kim, Kyle Min, Suhyeon Ha +2
Multi-perturbation adversarial training (MAT) aims to achieve robustness against multiple perturbations but suffers from robustness trade-offs between different threats. T…
Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
Chungpa Lee, Jihoon Kwon, Kyle Min +1
Vision-language models map images and text into a joint embedding space. However, these embeddings often entangle multiple semantic features, which limits their interpretability an…
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
Kibum Kim, Jiwan Kim, Kyle Min +4
Video Large Language Models (Video LLMs) incur high inference latency due to a large number of visual tokens provided to LLMs. To address this, training-free visual token pruning h…
FLAIR: Frequency- and Locality-Aware Implicit Neural Representations
Sukhun Ko, Seokhyun Youn, Dahyeon Kye +3
Implicit Neural Representations (INRs) leverage neural networks to map coordinates to corresponding signals, enabling continuous and compact representations. This paradigm has driv…
ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries
Junhyuk Kwon, Seungjoon Lee, Hyejin Park +2
Natural-language instance navigation becomes challenging when the initial user request does not uniquely specify the target instance. A practical agent should reduce the user's bur…
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
Julia Lee Romero, Kyle Min, Subarna Tripathi +1
Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We…