Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Benchmark Everything Everywhere All at Once
Shiyun Xiong, Dongming Wu, Peiwen Sun +5
Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensiv…
cs.AI2026
Towards Autonomous Business Intelligence via Data-to-Insight Discovery Agent
Dongming Wu, Junwen Li, Ming Lu +2
Transforming fragmented enterprise data into actionable insights remains a significant challenge for LLMs, constrained by complex database schemas, limitations in dynamic SQL gener…