Hidden in Plain Text: Emergence and Mitigation of Steganographic Collusion in LLMs
Paper accepted to Safe Generative AI Workshop at NeurIPS 2024
Steganography is the practice of concealing information within a message while avoiding detection by a third party. There is a risk that future multi-agent LLM systems may learn to collude via the use of text-based steganography, allowing them to execute misaligned behaviour without being detected by an overseer. In this paper, we investigate the emergence of steganography and investigate the effectiveness of various mitigation strategies. First, we design a set of techniques for eliciting steganography in LLMs – a gradient-based reinforcement learning (GBRL) and an in-context reinforcement learning (ICRL) method. Second, we leverage these elicitation methods to demonstrate that steganography can indeed arise as an instrumental goal whilst optimizing task performance in some settings. Third, we use our ICRL elicitation method to generate strategies that can be used to red-team steganography mitigation techniques. Here, we find that oversight of model outputs alone is insufficient for preventing steganography. Additionally, while combining paraphrasing of model outputs with oversight can help, this is insufficient for mitigating 1-bit steganography. Overall, these results motivate the need for further research into steganography as a phenomenon in LLMs, and for the development of improved steganography mitigation methods.
Programme
LASR
LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.