12 papers
HiTMS: A High-Throughput Multi-Stream Linguistic Steganography Framework
Ruiyi Yan, Yugo Murawaki, Zhongliang Yang
Generative linguistic steganography conceals secret bits within the sampling randomness of large language models. Existing schemes are single-stream, conveying an entire secret thr…
Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition
Shiho Matta, Yin Jou Huang, Fei Cheng +3
We introduce Dango, a 1.8B-parameter large language model designed for controlled studies of L1-to-L2 (Japanese-to-English) transfer in second language acquisition (SLA). While pre…
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Keizo Kato, Chenhui Chu, Yugo Murawaki +1
For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable progress. But they typically rely on large…
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
Ruiyi Yan, Shiao Meng, Yugo Murawaki
Linguistic steganography based on language models typically assumes that steganographic texts are transmitted without alteration, making them fragile to even minor modifications. W…
Efficient Provably Secure Linguistic Steganography via Range Coding
Ruiyi Yan, Yugo Murawaki
Linguistic steganography involves embedding secret messages within seemingly innocuous texts to enable covert communication. Provable security, which is a long-standing goal and ke…
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models
Chengzhi Zhong, Fei Cheng, Qianying Liu +3
Large language models exhibit strong multilingual capabilities despite limited exposure to non-English data. Prior studies show that English-centric large language models map multi…