nerdexam
Amazon

AIF-C01 · Question #43

A social media company wants to use a large language model (LLM) for content moderation. The company wants to evaluate the LLM outputs for bias and potential discrimination against specific groups…

The correct answer is D. Benchmark datasets. Benchmark datasets are pre-validated datasets specifically designed to evaluate machine learning models for bias, fairness, and potential discrimination. These datasets are the most efficient tool for assessing an LLM's performance against known standards with minimal…

Submitted by lucia.co· Mar 30, 2026Model Evaluation

Question

A social media company wants to use a large language model (LLM) for content moderation. The company wants to evaluate the LLM outputs for bias and potential discrimination against specific groups or individuals. Which data source should the company use to evaluate the LLM outputs with the LEAST administrative effort?

Options

  • AUser-generated content
  • BModeration logs
  • CContent moderation guidelines
  • DBenchmark datasets

How the community answered

(27 responses)
  • A
    11% (3)
  • B
    4% (1)
  • C
    4% (1)
  • D
    81% (22)

Explanation

Benchmark datasets are pre-validated datasets specifically designed to evaluate machine learning models for bias, fairness, and potential discrimination. These datasets are the most efficient tool for assessing an LLM's performance against known standards with minimal administrative effort.

Topics

#LLM evaluation#Bias detection#Content moderation#Benchmark datasets

Community Discussion

No community discussion yet for this question.

Full AIF-C01 Practice