bilingual rendering 1image editing 1multimodal understanding 1open-source models 1text-to-image generation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Guoxuan Chen, Chufeng Xiao, Haoran Yang +30
Boogu-Image-0.1 is an open-source multimodal model family that supports high-quality text-to-image generation, fast inference, instruction-based image editing, and bilingual (Chine…
cs.CV2026
PartHOI: Part-based Hand-Object Interaction Transfer via Generalized Cylinders
Qiaochu Wang, Chufeng Xiao, Manfred Lau +1
Learning-based methods to understand and model hand-object interactions (HOI) require a large amount of high-quality HOI data. One way to create HOI data is to transfer hand poses…
cs.HC2025
MoGraphGPT: Creating Interactive Scenes Using Modular LLM and Graphical Control
Hui Ye, Chufeng Xiao, Jiaye Leng +2
Creating interactive scenes often involves complex programming tasks. Although large language models (LLMs) like ChatGPT can generate code from natural language, their output is of…