Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
Yiming Huang, Ziche Liu, Zhuohang Wu +11
Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hypothesis generation, observational testing, and refinement unde…
cs.AI2026
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
Han Wang, Yifan Sun, Brian Ko +8
Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer…
cs.AI2026
How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design Factors
Kuai Yu, Naicheng Yu, Han Wang +2
Web agents have demonstrated strong performance on a wide range of web-based tasks. However, existing research on the effect of environmental variation has mostly focused on robust…