6 papers
TIEM: Temporal Integration of Hypergraph Evidence and Skill Memory for Event-Driven Financial Forecasting
Wenjin Liu, Shen Pang, Chenxi Wang +5
Event-driven catalyst-outcome forecasting increasingly uses retrieval- and memory-augmented large language model agents for prediction. However, training-data contamination and tem…
FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing
Mingda Zhang, Wenjin Liu, Tiesunlong Shen +5
In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key challenges, including human-depend…
SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration
Mingda Zhang, Tiesunlong Shen, Haoran Luo +4
In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existing orchestration methods sti…
LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence
Wenjin Liu, Haoran Luo, Xin Feng +6
Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulating the expertise of legal exper…
Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
Wenjin Liu, Haoran Luo, Xueyuan Lin +5
Recently, advanced large language models (LLMs) have emerged at an increasingly rapid pace. However, when faced with complex problems, most users are often unable to provide accura…
BenchBench: Benchmarking Automated Benchmark Generation
Yandan Zheng, Haoran Luo, Zhenghong Lin +2
Benchmarks are the de facto standard for tracking progress in large language models (LLMs), yet static test sets can rapidly saturate, become vulnerable to contamination, and are c…