DEA-C01 · Question #142
A company uploads .csv files to an Amazon S3 bucket. The company's data platform team has set up an AWS Glue crawler to perform data discovery and to create the tables and schemas. An AWS Glue job wri
The correct answer is A. Modify the AWS Glue job to copy the rows into a staging Redshift table. Add SQL commands to. Two step approach involving creating a staging table, followed by using Redshift's merge statement to update the target table from staging table and finally truncate/housekeep the staging
Question
A company uploads .csv files to an Amazon S3 bucket. The company's data platform team has set up an AWS Glue crawler to perform data discovery and to create the tables and schemas. An AWS Glue job writes processed data from the tables to an Amazon Redshift database. The AWS Glue job handles column mapping and creates the Amazon Redshift tables in the Redshift database appropriately. If the company reruns the AWS Glue job for any reason, duplicate records are introduced into the Amazon Redshift tables. The company needs a solution that will update the Redshift tables without duplicates. Which solution will meet these requirements?
Options
- AModify the AWS Glue job to copy the rows into a staging Redshift table. Add SQL commands to
- BModify the AWS Glue job to load the previously inserted data into a MySQL database. Perform an
- CUse Apache Spark's DataFrame dropDuplicates() API to eliminate duplicates. Write the data to
- DUse the AWS Glue ResolveChoice built-in transform to select the value of the column from the
How the community answered
(43 responses)- A79% (34)
- B5% (2)
- C14% (6)
- D2% (1)
Explanation
Two step approach involving creating a staging table, followed by using Redshift's merge statement to update the target table from staging table and finally truncate/housekeep the staging
Topics
Community Discussion
No community discussion yet for this question.