3 papers
cs.CV2026
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
Svetlana Orlova, Niccolò Cavagnero, Gijs Dubbelman
Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive video datasets, resulting in sub…
cs.CV2025
Simplifying Traffic Anomaly Detection with Video Foundation Models
Svetlana Orlova, Tommie Kerssies, Brunó B. Englert +1
Recent methods for ego-centric Traffic Anomaly Detection (TAD) often rely on complex multi-stage or multi-representation fusion architectures, yet it remains unclear whether such c…
cs.CV2024
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
Narges Norouzi, Svetlana Orlova, Daan de Geus +1
This work presents Adaptive Local-then-Global Merging (ALGM), a token reduction method for semantic segmentation networks that use plain Vision Transformers. ALGM merges tokens in…