8 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
Layered Division and Global Allocation for Community Detection in Multilayer Network
Fanghao Hu, Zhi Cai, Bang Wang
Community detection in multilayer networks (CDMN) is to divide a set of entities with multiple relation types into a few disjoint subsets, which has many applications in the Web, t…
VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction
Junhong Lin, Kangli Wang, Shunzhou Wang +3
Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while ele…
HeroFilter: Adaptive Spectral Graph Filter for Varying Heterophilic Relations
Shuaicheng Zhang, Haohui Wang, Junhong Lin +5
Graph heterophily, where connected nodes have different labels, has attracted significant interest recently. Most existing works adopt a simplified approach - using low-pass filter…
Temporal Reasoning with Large Language Models Augmented by Evolving Knowledge Graphs
Junhong Lin, Song Wang, Xiaojie Guo +2
Large language models (LLMs) excel at many language understanding tasks but struggle to reason over knowledge that evolves. To address this, recent work has explored augmenting LLM…
LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
Xinyue Zeng, Haohui Wang, Junhong Lin +3
The proliferation of open-sourced Large Language Models (LLMs) and diverse downstream tasks necessitates efficient model selection, given the impracticality of fine-tuning all cand…