1 citations · 3 across the 4 of their papers we have counts for
4 papers
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Fei Wang, Xingyu Fu, James Y. Huang +18
We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tas…
MIRAI: Evaluating LLM Agents for Event Forecasting
Chenchen Ye, Ziniu Hu, Yihe Deng +4
Recent advancements in Large Language Models (LLMs) have empowered LLM agents to autonomously collect world information, over which to conduct reasoning to solve complex problems.…
Improving Event Definition Following For Zero-Shot Event Detection
Zefan Cai, Po-Nien Kung, Ashima Suvarna +6
Existing approaches on zero-shot event detection usually train models on datasets annotated with known event types, and prompt them with unseen event definitions. These approaches…
Instructional Fingerprinting of Large Language Models
Jiashu Xu, Fei Wang, Mingyu Derek Ma +3
The exorbitant cost of training Large language models (LLMs) from scratch makes it essential to fingerprint the models to protect intellectual property via ownership authentication…