Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing
Shaoan Zhao, Fang Zhao, Xueqiang Guo +9
Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into ques…
cs.AI2026
Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis
Yantao Li, Huanlin Gao, Fang Zhao +12
Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless g…
cs.AI2026
MediaClaw: Multimodal Intelligent-Agent Platform Technical Report
Shaoan Zhao, Huanlin Gao, Qiang Hui +9
MediaClaw is a multimodal agent platform built on the OpenClaw ecosystem. Its core design follows a three-layer architecture of unified abstraction, pluginized extension, and workf…