Atwood's original article was puzzling to me, and the conclusions just didn't compute.
I can recall at least a half dozen times when I was a DBA in olden times that ECC either corrected was essential in the isolation of faults on my Informix and later Oracle boxes, running mostly on Sun and RS/6000 at the time.
Sun had a nice habit of shipping defective CPUs and memory in the late 90s. The details are foggy, but I remember correlating ECC faults to long transactions that would fail, and getting a bunch of stuff out of Sun.
Than again, that was 15+ years ago, so maybe the newfangled memory we have these days is more reliable.
Here's James Gosling's account of radioactive RAM chips used in the UltraSparc II...
http://nighthacks.com/roller/jag/entry/at_the_mercy_of_suppl...
When Sun folks get together and bullshit about their theories of why Sun died, the one that comes up most often is another one of these supplier disasters. Towards the end of the DotCom bubble, we introduced the UltraSPARC-II. Total killer product for large datacenters. We sold lots. But then reports started coming in of odd failures. Systems would crash strangely. We'd get crashes in applications. All applications. Crashes in the kernel. Not very often, but often enough to be problems for customers. Sun customers were used to uptimes of years. The US-II was giving uptimes of weeks. We couldn't even figure out if it was a hardware problem or a software problem - Solaris had to be updated for the new machine, so it could have been a kernel problem. But nothing was reproducible. We'd get core dumps and spend hours pouring over them. Some were just crazy, showing values in registers that were simply impossible given the preceeding instructions. We tried everything. Replacing processor boards. Replacing backplanes. It was deeply random. It's very randomness suggested that maybe it was a physics problem: maybe it was alpha particles or cosmic rays. Maybe it was machines close to nuclear power plants. One site experiencing problems was near Fermilab. We actually mapped out failures geographically to see if they correlated to such particle sources. Nope. In desperation, a bright hardware engineer decided to measure the radioactivity of the systems themselves. Bingo! Particles! But from where? Much detailed scanning and it turned out that the packaging of the cache ram chips we were using was noticeably radioactive. We switched suppliers and the problem totally went away. After two years of tearing out hair out, we had a solution.
But it was too late. We had spent billions of dollars keeping our customers running. Swapping out all of that hardware was cripplingly expensive. But even worse, it severely damaged our customers trust in our products. Our biggest customers had been burned and were reluctant to buy again. It took quite a few years to rebuild that trust. At about the time that it felt like we had rebuilt trust and put the debacle behind us, the Financial Crisis hit...</i>
Interesting. I had not heard that story before, although I believe one or two of my old friends were working at Sun on Sparc development at the time. The curious thing about this story is that when I worked in the memory business for a while in the late '80s I was told that the major source of alpha particles that could cause soft errors was the device packaging material. As it was explained to me, earlier in DRAM history they were usually packaged in ceramic packages (also military applications always used ceramic). Later no plastic packaging materials were used because they were less expensive (and presumably the associated reliability issues had been resolved sufficiently to allow the use of plastic in more applications). Anyway, the plastic didn't emit alpha radiation like the ceramic did. When it was explained to me I got the impression everyone in the business knew this.
Don't hardware manufacturers test their systems for many weeks for MTBF estimates and the like? For something that is supposed to be running for years at a time, how did this escape their QA process before shipping?
We use high-spec systems for storing data coming off radiation detectors in experiments (which can be in the multiple GB/s of data with high end digitizers). You can bet we use ECC for that; we made sure to after one experiment got ruined by memory corruption...
Are you referring to the Ecache Data Parity (EDP) errors? I think those things hastened Sun's downfall. They replaced every single CPU and memory board in every Sun server when I was at Hotmail. That was a lot of modules.