3 papers
cs.CV2026
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
Ariel Shaulov, Eitan Shaar, Amit Edenzon +1
Auto-regressive video generation enables long video synthesis by iteratively conditioning each new batch of frames on previously generated content. However, recent work has shown t…
cs.CV2025
Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds
Eitan Shaar, Ariel Shaulov, Gal Chechik +1
In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing appr…
cs.CL2025
Classifier-Guided Captioning Across Modalities
Ariel Shaulov, Tal Shaharabany, Eitan Shaar +2
Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to gen…