AAISM · Question #85
When robust input controls are not practical on a large language model (LLM) to prevent prompt injection attacks from external threats, which of the following would be the BEST compensating control…
The correct answer is A. Review and annotate the AI system's outputs. When prompt injection cannot be reliably prevented at the input layer, a compensating control shifts focus to detecting and containing successful attacks at the output layer. Reviewing and annotating the LLM's outputs allows human or automated review to identify anomalous…
Question
When robust input controls are not practical on a large language model (LLM) to prevent prompt injection attacks from external threats, which of the following would be the BEST compensating control to address the risk?
Options
- AReview and annotate the AI system's outputs
- BImplement identity and access management (IAM)
- CConduct human reviews of the AI system's inputs
- DFine-tune the system to validate the AI system's inputs
How the community answered
(37 responses)- A76% (28)
- B8% (3)
- C14% (5)
- D3% (1)
Explanation
When prompt injection cannot be reliably prevented at the input layer, a compensating control shifts focus to detecting and containing successful attacks at the output layer. Reviewing and annotating the LLM's outputs allows human or automated review to identify anomalous, harmful, policy-violating, or adversarially influenced responses before they are delivered or acted upon - limiting the blast radius of a successful injection. IAM (B) manages access but doesn't address injected content within authorized sessions. Human review of inputs (C) is costly, slow, and impractical at scale for LLMs. Fine-tuning to validate inputs (D) is essentially an input control, not a compensating control when input controls are explicitly deemed impractical.
Topics
Community Discussion
No community discussion yet for this question.