3 papers
cs.CV2026
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
Haichen He, Jiayi Zhou, Sifeng Shang +3
Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video du…
cs.CV2025
Measuring Epistemic Humility in Multimodal Large Language Models
Bingkui Tong, Jiaer Xia, Sifeng Shang +1
Hallucinations in multimodal large language models (MLLMs) -- where the model generates content inconsistent with the input image -- pose significant risks in real-world applicatio…
cs.LG2025
Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
Sifeng Shang, Jiayi Zhou, Chenyu Lin +2
As the size of large language models grows exponentially, GPU memory has become a bottleneck for adapting these models to downstream tasks. In this paper, we aim to push the limits…