New Security Challenges for Large Language Models
A team of researchers from Intel, Idaho State University, and the University of Illinois has introduced new techniques to breach security filters in large language models (LLMs) like ChatGPT and Gemini. According to 404 Media, this could have serious implications for security.
In their study, they found that chatbots could be manipulated into providing forbidden information if queries are made in complicated or ambiguous ways or by citing non-existent sources. This approach is referred to as "information overload".
The researchers utilized a tool named InfoFlood, which automates the "overloading" process with information, resulting in disorientation of the systems and potentially allowing access to prohibited content that is typically blocked.
The vulnerability arises from the fact that models focus on the superficial structure of the text, failing to recognize harmful content hidden within. This opens doors for malicious actors to circumvent restrictions.
As part of responsible disclosure, the authors of the study plan to share their findings with companies working with large LLMs to enhance their security systems. They will also provide solutions to the problems they identified during their research.
"LLM models mainly rely on input and output protections to detect harmful content. InfoFlood can be used to train these protective mechanisms, allowing relevant information to be extracted from potentially dangerous queries, making the models more resilient to such attacks," the study states.



