2 papers
cs.CL2026
CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
Baoliang Tian, Yuxuan Si, Jilong Wang +13
Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largel…
cs.IR2026
MALLOC: Benchmarking the Memory-aware Long Sequence Compression for Large Sequential Recommendation
Qihang Yu, Kairui Fu, Zhaocheng Du +10
The scaling law, which indicates that model performance improves with increasing dataset and model capacity, has fueled a growing trend in expanding recommendation models in both i…