3 papers
cs.CL2025
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
Hongjin Su, Howard Yen, Mengzhou Xia +12
Existing retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually…
cs.LG2025
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
Lejla Skelic, Yan Xu, Matthew Cox +3
The role of Large Language Models (LLMs) has not been extensively explored in analog circuit design, which could benefit from a reasoning-based approach that transcends traditional…
cs.LG2025
Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
Hongjin Su, Ruoxi Sun, Jinsung Yoon +3
Autonomous agents powered by large language models (LLMs) have the potential to enhance human capabilities, assisting with digital tasks from sending emails to performing data anal…