collaborators

12 papers

cs.CL2026

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

Hongyu Luo, He Wang, Huihao Jing +6

Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an an…

cs.CL2026

NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems

Jiayu Liu, Rui Wang, Qing Zong +9

Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is wid…

cs.CL2026

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

Wei Fan, Yining Zhou, Mufan Zhang +8

While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint that applicable law must matc…

cs.SE2026

WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality

Chunyang Li, Yilun Zheng, Xinting Huang +5

The paradigm of LLM-as-a-judge is emerging as a scalable and efficient alternative to human evaluation, demonstrating strong performance on well-defined tasks. However, its reliabi…

cs.CL2025

The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning

Tianshi Zheng, Yixiang Chen, Chengxi Li +7

Chain-of-Thought (CoT) prompting has been widely recognized for its ability to enhance reasoning capabilities in large language models (LLMs). However, our study reveals a surprisi…

cs.CL2025

CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?

Qing Zong, Jiayu Liu, Tianshi Zheng +7

Accurate confidence calibration in Large Language Models (LLMs) is critical for safe use in high-stakes domains, where clear verbalized confidence enhances user trust. Traditional…