4 papers
Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
Seunghyuk Cho, Zhenyue Qin, Yang Liu +3
Plane geometry problem solving (PGPS) has recently gained significant attention as a benchmark to assess the multi-modal reasoning capabilities of large vision-language models. Des…
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
Seunghyuk Cho, Zhenyue Qin, Yang Liu +3
We introduce GeoDANO, a geometric vision-language model (VLM) with a domain-agnostic vision encoder, for solving plane geometry problems. Although VLMs have been employed for solvi…
HandCraft: Anatomically Correct Restoration of Malformed Hands in Diffusion Generated Images
Zhenyue Qin, Yiqun Zhang, Yang Liu +1
Generative text-to-image models, such as Stable Diffusion, have demonstrated a remarkable ability to generate diverse, high-quality images. However, they are surprisingly inept whe…
Visual Prompting in LLMs for Enhancing Emotion Recognition
Qixuan Zhang, Zhifeng Wang, Dylan Zhang +5
Vision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing. Nonetheless, the potential of using visual prompts for em…