5 papers
Compression and Retrieval: Implicit Memory Retrieval for Video World Models
Zhan Peng, Jie Ma, Huiqiang Sun +6
Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories remains a critical challenge…
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model
Jinghan Wu, Jing Li, Ivor W. Tsang +1
Visual information helps resolve ambiguity in coreference resolution, leading to notable performance gains. However, existing Multi-modal Coreference Resolution (MCR) methods requi…
Uncover and Unlearn Nuisances: Agnostic Fully Test-Time Adaptation
Ponhvoan Srey, Yaxin Shi, Hangwei Qian +2
Fully Test-Time Adaptation (FTTA) addresses domain shifts without access to source data and training protocols of the pre-trained models. Traditional strategies that align source a…
Towards Harmless Rawlsian Fairness Regardless of Demographic Prior
Xuanqian Wang, Jing Li, Ivor W. Tsang +1
Due to privacy and security concerns, recent advancements in group fairness advocate for model training regardless of demographic information. However, most methods still require p…
Alpha and Prejudice: Improving -sized Worst-case Fairness via Intrinsic Reweighting
Jing Li, Yinghua Yao, Yuangang Pan +3
Worst-case fairness with off-the-shelf demographics achieves group parity by maximizing the model utility of the worst-off group. Nevertheless, demographic information is often una…