bilingual rendering 1image editing 1multimodal understanding 1open-source models 1text-to-image generation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Guoxuan Chen, Chufeng Xiao, Haoran Yang +30
Boogu-Image-0.1 is an open-source multimodal model family that supports high-quality text-to-image generation, fast inference, instruction-based image editing, and bilingual (Chine…
cs.CL2025
DiLA: Enhancing LLM Tool Learning with Differential Logic Layer
Yu Zhang, Hui-Ling Zhen, Zehua Pei +4
Considering the challenges faced by large language models (LLMs) in logical reasoning and planning, prior efforts have sought to augment LLMs with access to external solvers. While…
cs.CL2025
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
Shengyin Sun, Yiming Li, Xing Li +8
Test-time scaling has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs) by allocating additional computational resources durin…