nerdexam
CompTIA

DY0-001 · Question #38

Which of the following issues should a data scientist be most concerned about when generating a synthetic data set?

The correct answer is D. The data set not being representative of the population. If synthetic data don't accurately mirror the real-world distributions and relationships, any models trained on them will perform poorly in deployment. Representativeness is the critical concern when generating synthetic data.

Data Governance and Ethics

Question

Which of the following issues should a data scientist be most concerned about when generating a synthetic data set?

Options

  • AThe data set consuming too many resources
  • BThe data set having insufficient features
  • CThe data set having insufficient row observations
  • DThe data set not being representative of the population

How the community answered

(56 responses)
  • A
    4% (2)
  • B
    2% (1)
  • C
    2% (1)
  • D
    93% (52)

Explanation

If synthetic data don't accurately mirror the real-world distributions and relationships, any models trained on them will perform poorly in deployment. Representativeness is the critical concern when generating synthetic data.

Topics

#synthetic data#data representativeness#population bias#data quality

Community Discussion

No community discussion yet for this question.

Full DY0-001 Practice