From the 5 of 66 papers with an AI index.
17 citations
- Tsinghua UniversityCN25 papers
- Peking UniversityCN24 papers
- Institute of Modern PhysicsCN23 papers
- University of Chinese Academy of SciencesCN23 papers
- Carnegie Mellon UniversityUS21 papers
- Istituto Nazionale di Fisica Nucleare, Laboratori Nazionali di FrascatiIT21 papers
- Istituto Nazionale di Fisica Nucleare, Sezione di PerugiaIT21 papers
- Nanjing Normal UniversityCN21 papers
- National Centre for Nuclear ResearchPL21 papers
- South China Normal UniversityCN21 papers
- Università degli Studi del Piemonte Orientale “Amedeo Avogadro”IT21 papers
- University of BristolGB21 papers
4 papers · 1 filter
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Bo Deng, Kang Zhou, Lifan Guo +6
Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cov…
Action-Aware Generative Sequence Modeling for Short Video Recommendation
Wenhao Li, Zihan Lin, Zhengxiao Guo +7
The paper proposes a new recommendation model, A2Gen, that treats user actions on short videos as temporal sequences and uses attention and hierarchical encoding to predict future…
SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection
Zhiyong Cao, Dunqiang Liu, Qi Dai +9
Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific business outcomes, such as acquiring social-med…
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
Zenghui Zhou, Man Li, Xiaoke Fang +3
Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing evaluations rely on static benchmarks, wh…