21 papers
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding
Yijia Lei, Jinzhao Li, Yichi Zhang +3
We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models…
MESA: Improving MoE Safety Alignment via Decentralized Expertise
Yitong Sun, Yao Huang, Teng Li +5
Mixture-of-Experts (MoE) architectures scale Large Language Models (LLMs) efficiently, enabling greater capacity with reduced computational cost by dynamically routing inputs to re…
IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams
Jinzhao Li, Yinuo Chen, Wenxuan Song +5
Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require proactive reasoning over cont…
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
Hanyu Li, Yichi Zhang, Speed Zhu +3
Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on end-to-end tasks such as issue…
Mixture of Complementary Agents for Robust LLM Ensemble
Yichi Zhang, Kevin Lu, Yuang Zhang +3
Multi-AI collaboration, such as ensembling or debating large language models (LLMs), is a promising paradigm for aggregating information and boosting performance. A foundational st…
Unveiling the Basin-Like Loss Landscape in Large Language Models
Huanran Chen, Yinpeng Dong, Zeming Wei +4
We discover the emergence of \textit{basins} in the loss landscape of large language models. As model scale increases, LLMs become progressively more resilient to random perturbati…