2 papers
cs.CV2026
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
Seokmin Lee, Yunghee Lee, Byeonghyun Pak +1
For robotic agents operating in dynamic environments, learning visual state representations from streaming video observations is essential for sequential decision making. Recent se…
cs.CV2025
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
Taein Son, Soo Won Seo, Jisong Kim +2
Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues,…