most citedCovert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.IR2026

Synthetic Data from Cross-Domain Events for Large-Scale Recommendation Systems

Xiangyu Wang, Yawen He, Shivendra Pratap Singh +12

Large-scale recommendation systems operate across diverse domains, yet they face the challenges of data sparsity and noisy implicit feedback. Traditional approaches mitigate this v…

cs.AI2026

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs

Bingjun Luo, Tony Wang, Chaoqi Chen +1

Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual tokens required. To improve efficie…

cs.AI2026

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

Bingjun Luo, Tony Wang, Hanqi Chen +1

Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficiently compressing visual tokens wh…

cs.LG2024

Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach

Tony T. Wang, John Hughes, Henry Sleight +7

Defending large language models against jailbreaks so that they never engage in a broadly-defined set of forbidden behaviors is an open problem. In this paper, we investigate the d…

cs.CR20241 cited

Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

Danny Halawi, Alexander Wei, Eric Wallace +3

Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety…