It's probably best to find a root cause for the first instance of a new class of problems. At first there was no way to be sure what part of the stack some of these were in. They ruled out their software, third-party libraries, and memory before focusing on a CPU and then which core of the CPU. Knowing that sometimes a single core goes bad means you can't randomly assign one test run to a CPU with no core affinity and reliably detect it. Yes, the fix is to disable that core or core complex if you have that kind of resolution, or to ditch the CPU if you must, or to pull the whole server if your scale demands it. But knowing it's a hardware error in the CPU and not chasing it through layers of code every time is invaluable.