1 citations · 1 across the 5 of their papers we have counts for
4 papers · 1 filter
Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
Wenzhen Luo, Wei Guan, Yifan Yao +6
We introduce Falcon, a cross-domain Chinese text-to-SQL benchmark grounded in an enterprise-compatible dialect (MaxCompute/Hive). It contains 600 Chinese questions over 28 database…
SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine
Xiaochen Wang, Junqing He, Liang Chen +5
Large Language Models with chain-of-thought prompting, such as OpenAI-o1, have shown impressive capabilities in natural language inference tasks. However, Multi-hop Question Answer…
Athena: Retrieval-augmented Legal Judgment Prediction with Large Language Models
Xiao Peng, Liang Chen
Recently, large language models (LLMs) like ChatGPT, LLaMA, and Claude have prevailed in countless domains, including legal scenarios. With LLMs' rapid technological progress, the…
Meta Semantic Template for Evaluation of Large Language Models
Yachuan Liu, Liang Chen, Jindong Wang +2
Do large language models (LLMs) genuinely understand the semantics of the language, or just memorize the training data? The recent concern on potential data contamination of LLMs h…