6 papers
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
Huanxi Liu, Kun Hu, Jiaqi Liao +6
The paper introduces MCPEvol-Bench, a benchmark that tests how well large language model agents adapt to changing tool interfaces and functionalities in Model Context Protocol (MCP…
Learning Fine-Grained Geometry for Sparse-View Splatting via Cascade Depth Loss
Wenjun Lu, Haodong Chen, Anqi Yi +4
Novel view synthesis is a fundamental task in 3D computer vision that aims to reconstruct photorealistic images from novel viewpoints given a set of posed images. However, reconstr…
DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting
Penghui Wen, Mengwei He, Patrick Filippi +5
Accurate short-term precipitation forecasting is critical for weather-sensitive decision-making in agriculture, transportation, and disaster response. Existing deep learning approa…
Charting Empirical Laws for LLM Fine-Tuning in Scientific Multi-Discipline Learning
Lintao Wang, Zhuqiang Lu, Yilin Zhu +6
While large language models (LLMs) have achieved strong performance through fine-tuning within individual scientific domains, their learning dynamics in multi-disciplinary contexts…
Benchmark^2: Systematic Evaluation of LLM Benchmarks
Qi Qian, Chengsong Huang, Jingwen Xu +13
The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose B…
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
Zhuqiang Lu, Zhenfei Yin, Mengwei He +4
Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual conten…