Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
Chandler Smith, Marwa Abdulhai, Manfred Diaz +83
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with bo…
cs.AI2025
DSBC : Data Science task Benchmarking with Context engineering
Ram Mohan Rao Kadiyala, Siddhant Gupta, Jebish Purbey +4
Recent advances in large language models (LLMs) have significantly impacted data science workflows, giving rise to specialized data science agents designed to automate analytical t…