3 papers
cs.CV2024
Detailed Object Description with Controllable Dimensions
Xinran Wang, Haiwen Zhang, Baoteng Li +5
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models(MLLM…
cs.CV2024
Disentangle and denoise: Tackling context misalignment for video moment retrieval
Kaijing Ma, Han Fang, Xianghao Zang +7
Video Moment Retrieval, which aims to locate in-context video moments according to a natural language query, is an essential task for cross-modal grounding. Existing methods focus…
cs.CL2024
Towards Robustness and Diversity: Continual Learning in Dialog Generation with Text-Mixup and Batch Nuclear-Norm Maximization
Zihan Wang, Jiayu Xiao, Mengxiang Li +4
In our dynamic world where data arrives in a continuous stream, continual learning enables us to incrementally add new tasks/domains without the need to retrain from scratch. A maj…