3 papers
cs.CV2026
Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought
Yu Huo, Siyu Zhang, Kun Zeng +7
Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints, notably generative numeracy…
cs.SI2026
A large-scale analysis of public-facing, community-built chatbots on Character.AI
Owen Lee, Kenneth Joseph
This paper presents the first large-scale analysis of public-facing chatbots on CharacterAI, a rapidly growing social media platform where users create and interact with chatbot…
cs.SD2025
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
Shunian Chen, Xinyuan Xie, Zheshu Chen +5
High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and con…