Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
Jianxin Gao, Beini Hu, Runze Li +6
With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks eval…
cs.AI2026
REED: Post-Training Representation Editing for Cross-Domain Linguistic Steganalysis
Ruohan Lei, Jianxin Gao, Wanli Peng +1
In real-world scenarios of linguistic steganalysis, tested texts usually come from unseen domains with different vocabularies, topics, writing styles, and steganographic generation…