AIGP · Question #103
When assessing the success of an AI system in meeting its objectives, which of the following approaches best aligns with the requirement to ensure a comprehensive evaluation while avoiding automation
The correct answer is D. Conduct a thorough review that includes the AI system's input and output data, human. Option D is correct because a thorough review combining both input/output data analysis and human judgment guards against automation bias - the tendency to uncritically accept AI outputs as correct simply because a machine produced them. Relying on diverse evaluation layers (data
Question
When assessing the success of an AI system in meeting its objectives, which of the following approaches best aligns with the requirement to ensure a comprehensive evaluation while avoiding automation bias?
Options
- AEvaluate and align pre-defined benchmarks which will provide evidence of having achieved your
- BRely on the AI system's output to determine if the goals were achieved as this will provide a
- CFocus on user feedback about the AI system's performance to directly measure the system's
- DConduct a thorough review that includes the AI system's input and output data, human
How the community answered
(50 responses)- A12% (6)
- B6% (3)
- C26% (13)
- D56% (28)
Explanation
Option D is correct because a thorough review combining both input/output data analysis and human judgment guards against automation bias - the tendency to uncritically accept AI outputs as correct simply because a machine produced them. Relying on diverse evaluation layers (data audit + human oversight) ensures no single mechanism dominates the assessment.
Why the distractors fail:
- A (pre-defined benchmarks only) is incomplete - benchmarks measure narrow, pre-specified behaviors and can be "gamed" or miss real-world failure modes outside their scope.
- B (rely on the AI's own output) is the textbook definition of automation bias - you're letting the system grade itself, which cannot surface systematic errors or misaligned objectives.
- C (user feedback only) captures perception but not ground truth; users may not notice subtle failures, and subjective satisfaction doesn't confirm the system actually met its objectives.
Memory tip: Think of the acronym HALO - Human + Automated + Layers + Oversight. Whenever a question mentions avoiding automation bias, the right answer will always include a human review component alongside data analysis, never automation alone.
Topics
Community Discussion
No community discussion yet for this question.