4 papers
LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models
Weibin Liao, Xin Gao, Tianyu Jia +6
Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive access to databases. While recent ap…
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
Xinrong Chen, Xu Chu, Yingmin Qiu +8
Large Vision-Language Models (LVLMs) can reason from image-text inputs and perform well in various multimodal tasks. Despite this success, they are affected by language priors and…
NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
X. Feng, H. Yu, M. Wu +6
With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation…
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
Meng-Hao Guo, Xuanyu Chu, Qianrui Yang +12
The rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini, and o3, with their capability to process and generate content across modalities s…