8 papers
GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation
Sifan Li, Yujun Cai, Hongkai Chen +1
Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-purpose vector code generation…
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
Liyang Chen, Hongkai Chen, Yujun Cai +3
Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remai…
OptiSQL: Executable SQL Generation from Optical Tokens
Sifan Li, Hongkai Chen, Yujun Cai +3
Executable SQL generation is typically studied in text-to-SQL settings, where tables are provided as fully linearized textual schemas and contents. While effective, this formulatio…
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
Liyang Chen, Hongkai Chen, Yujun Cai +3
Video-to-Audio generation has made remarkable strides in automatically synthesizing sound for video. However, existing evaluation metrics, which focus on semantic and temporal alig…
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
Sifan Li, Yujun Cai, Yiwei Wang
Vision-language models (VLMs) excel in semantic tasks but falter at a core human capability: detecting hidden content in optical illusions or AI-generated images through perceptual…
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
Sifan Li, Hongkai Chen, Yujun Cai +4
Vision Language Models (VLMs) have achieved impressive progress in multimodal reasoning; yet, they remain vulnerable to hallucinations, where outputs are not grounded in visual evi…