Prompt injection: AI instructions hidden in briefs
Prompt injection is a technique in which secret instructions are embedded in text to influence AI systems. A plaintiff in the US used this method to avoid possible interference with AI systems. Practice raises questions about the security and integrity of AI models, as unexpected input can lead to unwanted output. Such attacks are particularly relevant in the legal and scientific context where AI is used for the analysis and evaluation of texts. One startup, for example, used AI to evaluate scientific papers, leading to debates about transparency and bias. At the same time, it shows that AI systems are also used in companies to optimize processes, such as OpenAI, which has established a direct line to the boardroom to reduce bureaucracy. These developments underscore the importance of security measures and control mechanisms to protect AI systems from unauthorized manipulation. A central aspect is the clarification of legal framework conditions, as discussed in the AI Act. The webinar “AI Act Compact” highlights the implementation of the requirements that companies have in terms of labelling, AI competence and high-risk AI. The challenges of AI verification and transparency are therefore not only technical, but also legal in nature. In this context, it is critical that companies and organizations take action to protect AI systems from unauthorized input while ensuring the efficiency of their processes. Ensuring the integrity of AI systems is therefore a central focus in the development and application of modern AI technologies.