12 papers
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Junzhi Chen, Harsh Trivedi, Jane Pan +4
Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification que…
Certificates without Electrons? Theory and Evidence on Impacts from AI-Driven Power Demand
Dana Golden, Aruna Balasubramanian, Niranjan Balasubramanian
Data centers now account for 4.4% of United States electricity demand, yet the grid-level effectiveness of the renewable energy certificates (RECs) and power purchase agreements (P…
Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks
Dikshya Mohanty, Mohammad Saqib Hasan, Syed Mostofa Monsur +3
Research in AI4Science has shown promise in many science applications, including polymer design. However, current LLMs are ineffective in this problem space because: (i) most model…
Syntax Is Easy, Semantics Is Hard: Evaluating LLMs for LTL Translation
Priscilla Kyei Danso, Mohammad Saqib Hasan, Niranjan Balasubramanian +1
Propositional Linear Temporal Logic (LTL) is a popular formalism for specifying desirable requirements and security and privacy policies for software, networks, and systems. Yet ex…
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
Anurag Dutt, Young Won Choi, Avirup Sil +3
With the widespread adoption of Large Language Models (LLMs), energy costs of running LLMs is quickly becoming a critical concern. However, precisely measuring the energy consumpti…
ProST: Progressive Sub-task Training for Pareto-Optimal Multi-agent Systems Using Small Language Models
Biddut Sarker Bijoy, Mohammad Saqib Hasan, Pegah Alipoormolabashi +3
Multi-agent systems with smaller language models (SLMs) present a viable alternative to single agent systems powered by large language models (LLMs) for addressing complex problems…