8 papers
Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks
Shunfan Zheng, Dongsheng Shi, Yue Li +3
Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administrators (DBAs). However, current…
Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach
Tianbao Jiang, Weicong Ni, Gerard de Melo +1
Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning task…
From Construction to Injection: Edit-Based Fingerprints for Large Language Models
Yue Li, Xin Yi, Dongsheng Shi +3
Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial misuse. In black-box deployment, verificati…
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models
Zetian Ouyang, Linlin Wang, Gerard de Melo +1
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by…
AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
Yue Li, Xin Yi, Dongsheng Shi +3
Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in large vision language models (LVLMs). However, vision-agnostic water…
Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
Yue Li, Xin Yi, Dongsheng Shi +3
With the increasing size of Large Vision-Language Models (LVLMs), network pruning techniques aimed at compressing models for deployment in resource-constrained environments have ga…