10 papers
PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking
Dengxian Gong, Yuanzheng Wu, Haobo Yuan +11
This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leading to long, redundant trajectories. We attribute this failure t…
HunyuanImage 3.0 Technical Report
Tencent Hunyuan Foundation Model Team
We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module pub…
Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
Xinxin Liu, Zhaopan Xu, Ming Li +3
While Chain-of-Thought (CoT) prompting enables sophisticated symbolic reasoning in LLMs, it remains confined to discrete text and cannot simulate the continuous, physics-governed d…
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
HyperAI Team, Yuchen Liu, Kaiyang Han +26
Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy direc…
Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs
Yuren Mao, Peigen Liu, Xinjian Wang +11
Multi-Agent System (MAS) developing frameworks serve as the foundational infrastructure for social simulations powered by Large Language Models (LLMs). However, existing frameworks…
Explainable AI-Generated Image Detection RewardBench
Michael Yang, Shijian Deng, William T. Doan +4
Conventional, classification-based AI-generated image detection methods cannot explain why an image is considered real or AI-generated in a way a human expert would, which reduces…