I can reproduce this problem on my non-Threadripper Ryzen 5 3600 Zen 2 CPU. I don't think it's specific to TR.
With AVX2 enabled, Prime95's torture test is only stable when I use 3 workers or less. With 4 workers one of them will abort due to an error within 20 seconds. The more workers, the sooner a crash; with 5 workers it happens within 10 seconds, and with 6 workers it happens within 2-3 seconds.
If I play with the tests on and off for a while, seemingly increasing the quiescent temperature of the CPU, the whole experience and testing actually becomes a bit more stable. My motherboard uses the B450M chipset.
I cannot reproduce it on Ryzen 5 3600, Gigabyte B450M DS3H, Prime95 on Linux. (FMA3 FFT length 16K)
What motherboard make/model?
Gigabyte B450M DS3H. Currently running with latest BIOS (F50).
Note that @DerAlbi thinks the CPU is fine, instead he suspects the VRM on his Gigabyte mobo for this issue (based on his oscilloscope readings):
https://forum.level1techs.com/t/3970x-prime95-stability/1532...
Thanks for the info!
Do you have PBO or overlocking enabled? PBO is not considered stock by AMD, ensure that it's disabled in BIOS (some motherboards incorrectly enable it by default)
I don't do any manual overclocking, but "core performance boost" is enabled which allows the CPU to ramp up about 500 MHz during heavy load instead of staying fixed at base frequency. I'm guessing this is the same setting you're referring to?
Add.: OK, I finally found a couple of "Precision Boost Overdrive" menus deep down in the overclocking settings of this board's BIOS, and I've disabled everything PBO I can find while letting the "core performance boost" remain in auto mode. Without doing any extended testing this seems to have solved the problem, as I can now let Prime95 run through the AVX2 code path with 6 workers without any crashes. Thanks for the hint!