From the 1 of 55 linked papers with an AI index.
1 citations · 1 across the 4 of their papers we have counts for
55 papers
Experimental Designs for Multi-Item Multi-Period Inventory Control
Xinqi Chen, Xingyu Bai, Zeyu Zheng +1
The paper studies how to run A/B tests in multi‑item, multi‑period inventory systems, showing that common designs like switchbacks and item‑level randomization can be biased, and p…
3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
Zhenyu Zhao, Nanshan Jia, Jihyeon Je +7
Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the…
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Kaiwen Xiong, Haonian Ji, Shi Qiu +4
Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and…
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data
Xinrui Ruan, Zhenyu Zhao, Waverly Wei +4
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Orga…
VisualClaw: A Real-Time, Personalized Agent for the Physical World
Haoqin Tu, Jianwen Chen, Zijun Wang +14
Vision language models are serving as general-purpose interfaces for complex multimodal tasks. However, deployment still faces three gaps: VLMs typically incur high latency and cos…
Experimentation for Different Scheduling Policies on Queues: Mixed Differences-in-Q Estimators Based on Little's Law
Nanshan Jia, Ramesh Johari, Nian Si +1
In data centers, tasks are dispatched to various servers to evenly distribute the workload. When a data center considers implementing a new scheduling algorithm, it typically condu…