DAS-C01 · Question #39
A company uses Amazon Elasticsearch Service (Amazon ES) to store and analyze its website clickstream data. The company ingests 1 TB of data daily using Amazon Kinesis Data Firehose and stores one…
The correct answer is C. Decrease the number of Amazon ES shards for the index. The root cause of both the slow queries and JVMMemoryPressure errors is over-sharding. AWS best practice recommends keeping Elasticsearch/OpenSearch shard sizes between 10–50 GB. With 1 TB of data, 1,000 shards averages only 1 GB per shard - far too many. Each shard consumes…
Question
A company uses Amazon Elasticsearch Service (Amazon ES) to store and analyze its website clickstream data. The company ingests 1 TB of data daily using Amazon Kinesis Data Firehose and stores one day's worth of data in an Amazon ES cluster. The company has very slow query performance on the Amazon ES index and occasionally sees errors from Kinesis Data Firehose when attempting to write to the index. The Amazon ES cluster has 10 nodes running a single index and 3 dedicated master nodes. Each data node has 1.5 TB of Amazon EBS storage attached and the cluster is configured with 1,000 shards. Occasionally, JVMMemoryPressure errors are found in the cluster logs. Which solution will improve the performance of Amazon ES?
Options
- AIncrease the memory of the Amazon ES master nodes.
- BDecrease the number of Amazon ES data nodes.
- CDecrease the number of Amazon ES shards for the index.
- DIncrease the number of Amazon ES shards for the index.
How the community answered
(64 responses)- A3% (2)
- B6% (4)
- C73% (47)
- D17% (11)
Explanation
The root cause of both the slow queries and JVMMemoryPressure errors is over-sharding. AWS best practice recommends keeping Elasticsearch/OpenSearch shard sizes between 10–50 GB. With 1 TB of data, 1,000 shards averages only 1 GB per shard - far too many. Each shard consumes JVM heap memory for segment metadata and other overhead, regardless of how much data it contains. With 1,000 shards across 10 nodes (100 shards/node), JVM memory is exhausted, causing pressure errors and degrading query performance. Kinesis Firehose write errors are a downstream symptom of the cluster being overwhelmed. Reducing shards to a few dozen (e.g., 20–30 shards for 1 TB with a replica) would dramatically reduce JVM overhead, restore memory headroom, and improve both query performance and ingestion reliability.
Topics
Community Discussion
No community discussion yet for this question.