19 papers
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
Haochen Zhao, Yongxiu Xu, Xinkui Lin +6
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged…
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Kaixin Ma, Di Feng, Alexander Metz +3
The paper introduces MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents that handle multi-image, multi-turn tasks across hundreds of too…
Apple Intelligence Foundation Language Models
Tom Gunter, Zirui Wang, Chong Wang +152
We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficiently on devices and a large serv…
Conditionally Site-Independent Neural Evolution of Antibody Sequences
Stephen Zhewen Lu, Aakarsh Vermani, Kohei Sanno +4
Common deep learning approaches for antibody engineering focus on modeling the marginal distribution of sequences. By treating sequences as independent samples, however, these meth…
Learning Structure, Energy, and Dynamics: A Survey of Artificial Intelligence for Protein Dynamics
Haocheng Tang, Liang Shi, Ya-Shi Zhang +3
Protein dynamics underlie many biological functions, yet remain difficult to characterize due to the high computational cost of molecular dynamics simulations and the scarcity of d…
COMPASS: Benchmarking Constrained Optimization in LLM Agents
Tian Qin, Felix Bai, Ting-Yao Hu +8
Human decision-making often involves constrained optimization. As LLM agents are deployed to assist with real-world tasks like travel planning, shopping, and scheduling, they must…