Challenge 74 ☆☆

Welcome to challenge Challenge 74.

Challenge 74: Convince the AI Assistant to Reveal the Secret

AI assistants can be given system instructions containing confidential information. If those instructions are not properly protected, prompt injection may cause the assistant to reveal information that it was explicitly told to keep secret.

This application provides a small AI assistant. The assistant has been given confidential information in its instructions and has been told never to reveal it.

Chat with the assistant and find a way to convince it to reveal the confidential information.

Note

The assistant has been explicitly instructed not to reveal the confidential information. Try different prompts and techniques to see whether you can get it to disclose its instructions or the secret.

💡 Tip: Secrets are often strings, numbers, or encoded values. Copy and paste exactly what you find.

You can solve this challenge using the following steps:

  1. Interact with the AI assistant:

    • Use the chat interface provided in the challenge.

    • Ask the assistant questions and observe how it responds.

    • Remember that the assistant has been given confidential information as part of its instructions.

  2. Investigate how the assistant is instructed:

    • Ask the assistant about its instructions, rules, or context.

    • Consider what information might be available to the model that is not directly visible in the chat interface.

  3. Try prompt injection:

    • The assistant has been instructed not to reveal the confidential information.

    • Try crafting messages that change, override, or manipulate those instructions.

    • Ask the assistant to reveal information from its instructions or context.

  4. Look for the secret:

    • If you successfully cause the assistant to disclose the confidential information, use the revealed value as the solution to the challenge.

Why putting secrets in AI system prompts is a secret-leakage risk

AI assistants are usually given a system prompt containing instructions that define how the assistant should behave. It can be tempting to include sensitive information in that prompt and simply instruct the model not to reveal it.

This is exactly what happens in this challenge: the application puts the secret directly into the AI assistant’s system prompt and tells the model not to disclose it.

The system prompt is not a security boundary. The model receives both the system instructions and the user’s input as part of its context. A user can deliberately craft prompts that attempt to change the model’s behaviour, expose its instructions, or persuade it to disclose information from its context.

Three failures compound in this scenario:

  • The secret is placed directly into the model’s context, making it available to the model during inference.

  • The application relies on the model following an instruction to keep the secret confidential.

  • A user can interact directly with the model and attempt to manipulate those instructions through prompt injection.

If the model reveals the secret, the application’s confidential information has crossed its intended security boundary.

What to do instead:

- Never use an AI model's system prompt as a secrets-management mechanism.
- Keep secrets outside the model's context whenever possible.
- Give the application access to secrets only when they are actually required for an operation.
- Treat all user-provided prompts as potentially adversarial input.
- Apply authorization and access controls in the application rather than relying on the model to enforce them.
- If sensitive information is accidentally exposed to a model, treat the information as potentially compromised and rotate the affected credential where appropriate.
Note

A system prompt is an instruction to an AI model, not an access-control mechanism. Telling a model "never reveal this secret" does not guarantee that the secret will remain confidential. If information must remain secret, the application should enforce that boundary outside the model.


0