11 papers · 1 filter
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
Woo Jae Kim, Kyle Min, Suhyeon Ha +2
Multi-perturbation adversarial training (MAT) aims to achieve robustness against multiple perturbations but suffers from robustness trade-offs between different threats. T…
FLAIR: Frequency- and Locality-Aware Implicit Neural Representations
Sukhun Ko, Seokhyun Youn, Dahyeon Kye +3
Implicit Neural Representations (INRs) leverage neural networks to map coordinates to corresponding signals, enabling continuous and compact representations. This paradigm has driv…
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
Julia Lee Romero, Kyle Min, Subarna Tripathi +1
Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We…
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
Jihoon Kwon, Kyle Min, Jy-yong Sohn
Despite recent advances, vision-language models trained with standard contrastive objectives still struggle with compositional reasoning -- the ability to understand structured rel…
DecompDreamer: A Composition-Aware Curriculum for Structured 3D Asset Generation
Utkarsh Nath, Rajeev Goel, Rahul Khurana +5
Current text-to-3D methods excel at generating single objects but falter on compositional prompts. We argue this failure is fundamental to their optimization schedules, as simultan…
ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning
Jongseo Lee, Kyungho Bae, Kyle Min +2
In this work, we tackle the problem of video classincremental learning (VCIL). Many existing VCIL methods mitigate catastrophic forgetting by rehearsal training with a few temporal…