Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
Yinhao Tang, Youqing Fang, Yanan Sun +6
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retri…
cs.AI2026
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
Zerun Ma, Guoqiang Wang, Xinchen Xie +7
While Large Language Models (LLMs) have empowered AI research agents to perform isolated scientific tasks, automating complex, real-world workflows, such as LLM training, remains a…
cs.AI2025
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
Kexian Tang, Junyao Gao, Yanhong Zeng +6
Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning across multiple sequential step…