3 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2026
CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution
Xiangxi Zheng, Kuang He, Jiayi Hu +6
Chart-to-code generation demands strict visual precision and syntactic correctness from Vision-Language Models (VLMs). However, existing approaches are fundamentally constrained by…
cs.CV2024★ 1 cited
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
Chris Kelly, Luhui Hu, Jiayin Hu +7
The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…
cs.CV2024★ 3 cited
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
Chris Kelly, Luhui Hu, Bang Yang +7
With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achie…