collaborators

5 papers

cs.DB2026

CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs

Zihan Nan, Yang Gu, Wei Liu +4

Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primaril…

cs.AI2026

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows

Wei Liu, Yang Gu, Xi Yan +5

Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines. While recent LLM-based approac…

cs.LG2025

Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios

Yunkai Dang, Mengxi Gao, Yibo Yan +8

Multimodal large language models (MLLMs) have recently achieved state-of-the-art performance on tasks ranging from visual question answering to video understanding. However, existi…

cs.IR2025

HyperG: Hypergraph-Enhanced LLMs for Structured Knowledge

Sirui Huang, Hanqian Li, Yanggan Gu +3

Given that substantial amounts of domain-specific knowledge are stored in structured formats, such as web data organized through HTML, Large Language Models (LLMs) are expected to…

cs.CL2025

Capturing Nuanced Preferences: Preference-Aligned Distillation for Small Language Models

Yanggan Gu, Junzhuo Li, Sirui Huang +3

Aligning small language models (SLMs) with human values typically involves distilling preference knowledge from large language models (LLMs). However, existing distillation methods…