Post-Mortem: How a 2-Line Config Change Brought Down Our Search Cluster for 42 Minutes
At 14:02 UTC on a Tuesday, our search cluster stopped responding. Here is how an innocent thread pool config change killed node 3.
The Root Cause
We had modified search thread pool queue size from 1000 to unbounded in an attempt to prevent request rejection under traffic spikes.