4 papers
Rethinking Expressivity and Efficiency in Test-Time Training
Zeyun Zhong, Joya Chen, Manuel Martin +3
Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token updat…
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
David J. Lerch, Livien Majer, Zeyun Zhong +3
Robust self-supervised learning of multi-modal video representations is critical for real-world applications such as driver distraction detection, where multiple sensors provide co…
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
David J. Lerch, Sarath Mulugurthi, Manuel Martin +2
Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on general datasets and struggle…
FlowNar: Scalable Streaming Narration for Long-Form Videos
Zeyun Zhong, Manuel Martin, Chengzhi Wu +4
Recent Large Multimodal Models (LMMs), primarily designed for offline settings, are ill-suited for the dynamic requirements of streaming video. While recent online adaptations impr…