3 papers
cs.CV2026
Measuring Epistemic Humility in Multimodal Large Language Models
Bingkui Tong, Jiaer Xia, Sifeng Shang +1
Hallucinations in multimodal large language models (MLLMs) -- where the model generates content inconsistent with the input image -- pose significant risks in real-world applicatio…
cs.CV2026
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
Haichen He, Jiayi Zhou, Sifeng Shang +3
Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video du…
cs.LG2026
Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
Sifeng Shang, Jiayi Zhou, Chenyu Lin +2
As the size of large language models grows exponentially, GPU memory has become a bottleneck for adapting these models to downstream tasks. In this paper, we aim to push the limits…