2 papers
cs.CL2025
ChronusOmni: Improving Time Awareness of Omni Large Language Models
Yijing Chen, Yihan Wu, Kaisi Guan +4
Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target v…
cs.CV2024
See or Guess: Counterfactually Regularized Image Captioning
Qian Cao, Xu Chen, Ruihua Song +3
Image captioning, which generates natural language descriptions of the visual information in an image, is a crucial task in vision-language research. Previous models have typically…