activity
20242026
collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

Xin Sun, Di Wu, Yuchen Guo +4

LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidenc…

cs.AI2026

SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity

Yifan Mo, Xiao Fu, Yue Su +4

This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts. Prior work faces challenges in unstructured grounding,…

cs.AI2026

Herculean: An Agentic Benchmark for Financial Intelligence

Xueqing Peng, Zhuohan Xie, Yupeng Cao +60

As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…

cs.AI2025

Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment

Jiahuan Pei, Fanghua Ye, Xin Sun +3

Large language models (LLMs) have advanced virtual educators and learners, bridging NLP with AI4Education. Existing work often lacks scalability and fails to leverage diverse, larg…

cs.AI2025

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

Haochen Huang, Jiahuan Pei, Yue Su +11

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precis…