3 papers
cs.CV2025
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
Silin Gao, Sheryl Mathew, Li Mi +6
Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to…
cs.AI2024
Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail +6
Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new er…
cs.LG2023
Difference-Masking: Choosing What to Mask in Continued Pretraining
Alex Wilf, Syeda Nahida Akter, Leena Mathur +5
The self-supervised objective of masking-and-predicting has led to promising performance gains on a variety of downstream tasks. However, while most approaches randomly mask tokens…