1 citations · 2 across the 7 of their papers we have counts for
8 papers
Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
Yicheng Pan, Zhenrong Zhang, Pengfei Hu +6
Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. Howe…
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
Shuhang Liu, Zhenrong Zhang, Pengfei Hu +7
Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrec…
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
Pengfei Hu, Zhenrong Zhang, Qikai Chang +8
Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out…
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
Yusheng Dai, Chenxi Wang, Chang Li +7
This paper introduces Swap Forward (SaFa), a modality-agnostic and efficient method to generate seamless and coherence long spectrum and panorama through latent swap joint diffusio…
Skeleton and Font Generation Network for Zero-shot Chinese Character Generation
Mobai Xue, Jun Du, Zhenrong Zhang +5
Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation…
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
Haotian Wang, Yuzhe Weng, Yueyan Li +10
Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In t…