4 papers
HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution
Xiaotian Luo, Fengxingyu Wang, Chuanrui Hu +2
Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is governed by the surrounding agent…
SoMe: A Realistic Benchmark for LLM-based Social Media Agents
Dizhan Xue, Jing Cui, Shengsheng Qian +2
Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM…
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
Zhenyu Yang, Yuhang Hu, Zemin Du +6
Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicabi…
Short-video Propagation Influence Rating: A New Real-world Dataset and A New Large Graph Model
Dizhan Xue, Shengsheng Qian, Chuanrui Hu +1
Short-video platforms have gained immense popularity, captivating the interest of millions, if not billions, of users globally. Recently, researchers have highlighted the significa…