nerdexam
CompTIA

DY0-001 · Question #18

A data scientist is developing a model to predict the outcome of a vote for a national mascot. The choice is between tigers and lions. The full data set represents feedback from individuals…

The correct answer is D. Out-of-sample data. The aggregated feedback covers only 80% of respondents, mostly from a few professions and locations, so the model hasn't "seen" the remaining 20% (and those underrepresented groups). Its performance on those unseen subsets (out-of-sample data) is therefore the primary concern…

Modeling, Analysis, and Outcomes

Question

A data scientist is developing a model to predict the outcome of a vote for a national mascot. The choice is between tigers and lions. The full data set represents feedback from individuals representing 17 professions and 12 different locations. The following rank aggregation represents 80% of the data set:

Which of the following is the most likely concern about the model's ability to predict the outcome of the vote?

Exhibit

DY0-001 question #18 exhibit

Options

  • AInterpolated data
  • BExtrapolated data
  • CIn-sample data
  • DOut-of-sample data

How the community answered

(57 responses)
  • A
    7% (4)
  • B
    25% (14)
  • C
    11% (6)
  • D
    58% (33)

Explanation

The aggregated feedback covers only 80% of respondents, mostly from a few professions and locations, so the model hasn't "seen" the remaining 20% (and those underrepresented groups). Its performance on those unseen subsets (out-of-sample data) is therefore the primary concern for how well it will predict the actual vote.

Topics

#out-of-sample data#model generalization#prediction#sampling bias

Community Discussion

No community discussion yet for this question.

Full DY0-001 Practice