large language models 2ambiguity resolution 1benchmark dataset 1benchmarking 1causal reasoning 1code generation 1mechanistic reasoning 1personalized assistance 1scientific data analysis 1user history 1
From the 2 of 7 linked papers with an AI index.
2 citations · 3 across the 5 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
Zijian Xu, Wenshuo Zhang, Zisen Qin +4
The paper defines personalized ambiguity adaptation for coding assistants, introduces the CAPA benchmark to evaluate how well models use a user's past resolved sessions to handle r…
cs.AI2026
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
Chuhan Shi, Xiaoquan Ren, Sicheng Song +3
The paper presents SDABench, a capability-oriented benchmark that evaluates large language models on scientific data analysis tasks across biology, chemistry, environment, geograph…