4 papers
Dynamic Frequency Modulation for Controllable Text-driven Image Generation
Tiandong Shi, Ling Zhao, Ji Qi +2
The success of text-guided diffusion models has established a new image generation paradigm driven by the iterative refinement of text prompts. However, modifying the original text…
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
Zhidian Huang, Zijun Yao, Ji Qi +7
Recent advances in multimodal large language models (MLLMs) mark a shift from non-thinking models to post-trained reasoning models capable of solving complex problems through think…
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
V Team, Wenyi Hong, Wenmeng Yu +90
We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this…
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
Jianfa Chen, Emily Shen, Trupti Bavalatti +10
Robust content moderation classifiers are essential for the safety of Generative AI systems. In this task, differences between safe and unsafe inputs are often extremely subtle, ma…