2 papers
cs.CL2026
Phoenix-VL 1.5 Medium Technical Report
Team Phoenix, :, Arka Ray +29
We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Singapore context. Developed as a…
cs.CV2025
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
Yuhao Jia, Wenhan Tan
Diffusion-driven text-to-image (T2I) generation has achieved remarkable advancements in recent years. To further improve T2I models' capability in numerical and spatial reasoning,…