4 papers
TLRD: Teaching LLMs to Reason over Tabular Data with Tri-Level Rationale Distillation
Tianyuan Liang, Xuwei Tan, Lei Shi +6
Tabular data is a primary medium for storing real-world information, driving many industrial applications of machine learning. Traditional predictors achieve strong predictive perf…
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
Zhuchenyang Liu, Ziyu Hu, Yao Zhang +1
Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index storage overhead. Existing training-free…
Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs
Xuwei Tan, Ziyu Hu, Xueru Zhang
Machine learning models trained on real-world data often inherit and amplify biases against certain social groups, raising urgent concerns about their deployment at scale. While nu…
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
Ziyu Hu, Zhiqing Zhong, Weijian Zheng +6
The exponential growth of large language models has outpaced the capabilities of traditional CPU and GPU architectures due to the slowdown of Moore's Law. Dataflow AI accelerators…