3 papers
cs.CV2026
Exploring High-Order Self-Similarity for Video Understanding
Manjin Kim, Heeseung Kwon, Karteek Alahari +1
Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this wo…
cs.RO2025
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
Huiwon Jang, Sihyun Yu, Heeseung Kwon +3
Leveraging temporal context is crucial for success in partially observable robotic tasks. However, prior work in behavior cloning has demonstrated inconsistent performance gains wh…
cs.CV2025
Lightweight Structure-Aware Attention for Visual Understanding
Heeseung Kwon, Francisco M. Castro, Manuel J. Marin-Jimenez +2
Attention operator has been widely used as a basic brick in visual understanding since it provides some flexibility through its adjustable kernels. However, this operator suffers f…