Introducing a master/minion update system to work I ran a batch update to take a certain percentage out of the cluster.
Unfortunately I got my selection criteria wrong and pulled out all of one cluster and half of a second, halting a few thousand operations.
Luckily the monitoring system was very quick to alert me of this and using the same (wrong) selection criteria it was a fairly simple process to stop the update and put them all back in the cluster.
Takeaways?
The age old cliche of "With great power comes great responsibility".
Oh and have good monitoring!
Unfortunately I got my selection criteria wrong and pulled out all of one cluster and half of a second, halting a few thousand operations.
Luckily the monitoring system was very quick to alert me of this and using the same (wrong) selection criteria it was a fairly simple process to stop the update and put them all back in the cluster.
Takeaways? The age old cliche of "With great power comes great responsibility". Oh and have good monitoring!