6 papers
Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models
Yu Ma, Hongli Shi, Xinran Xu
Interactive video world models generate rollouts autoregressively under an action stream, yet they are trained and evaluated almost exclusively on factual prediction. We study coun…
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
Yu Ma, Hongli Shi, Jing Li +2
Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists i…
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
Li Zhang, Yuzhen Shi, Yiran Hu +15
Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order…
GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
Zhaofen Wu, Hanrong Zhang, Fulin Lin +9
To sustain coherent long-term interactions, Large Language Model (LLM) agents must navigate the tension between acquiring new information and retaining prior knowledge. Current uni…
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
Yuzhen Shi, Huanghai Liu, Yiran Hu +27
As large language models (LLMs) are increasingly applied to legal domain-specific tasks, evaluating their ability to perform legal work in real-world settings has become essential.…
Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions
Yiran Hu, Huanghai Liu, Chong Wang +15
Large language models (LLMs) are being increasingly integrated into legal applications, including judicial decision support, legal practice assistance, and public-facing legal serv…