Sizing Up Servers: Intel's Skylake-SP Xeon versus AMD's EPYC 7000 - The Server CPU Battle of the Decade?
by Johan De Gelas & Ian Cutress on July 11, 2017 12:15 PM EST- Posted in
- CPUs
- AMD
- Intel
- Xeon
- Enterprise
- Skylake
- Zen
- Naples
- Skylake-SP
- EPYC
SMT Integer Performance With SPEC CPU2006
Next, to test the performance impact of simultaneous multithreading (SMT) on a single core, we test with two threads on the same core. This way we can evaluate how well the core handles SMT.
Subtest | Application type | Xeon E5-2690 @ 3.8 | Xeon E5-2690 v3 @ 3.5 | Xeon E5-2699 v4 @ 3.6 | EPYC 7601 @3.2 | Xeon 8176 @ 3.8 |
400.perlbench | Spam filter | 39.8 | 43.9 | 47.2 | 40.6 | 55.2 |
401.bzip2 | Compression | 32.6 | 32.3 | 32.8 | 33.9 | 34.8 |
403.gcc | Compiling | 40.7 | 43.8 | 32.5 | 41.6 | 32.1 |
429.mcf | Vehicle scheduling | 44.7 | 51.3 | 55.8 | 44.2 | 56.6 |
445.gobmk | Game AI | 36.6 | 35.9 | 38.1 | 36.4 | 39.4 |
456.hmmer | Protein seq. analyses | 32.5 | 34.1 | 40.9 | 34.9 | 44.3 |
458.sjeng | Chess | 36.4 | 36.9 | 39.5 | 36 | 41.9 |
462.libquantum | Quantum sim | 75 | 73.4 | 89 | 89.2 | 91.7 |
464.h264ref | Video encoding | 52.4 | 58.2 | 58.5 | 56.1 | 75.3 |
471.omnetpp | Network sim | 25.4 | 30.4 | 48.5 | 26.6 | 42.1 |
473.astar | Pathfinding | 31.4 | 33.6 | 36.6 | 29 | 37.5 |
483.xalancbmk | XML processing | 43.7 | 53.7 | 78.2 | 37.8 | 78 |
Now on a percentage basis versus the single-threaded results, so that we can see how much performance we gained from enabling SMT:
Subtest | Application type | Xeon E5-2699 v4 @ 3.6 | EPYC 7601 @3.2 | Xeon 8176 @ 3.8 |
400.perlbench | Spam filter | 109% | 131% | 110% |
401.bzip2 | Compression | 137% | 141% | 128% |
403.gcc | Compiling | 137% | 119% | 131% |
429.mcf | Vehicle scheduling | 125% | 110% | 131% |
445.gobmk | Game AI | 125% | 150% | 127% |
456.hmmer | Protein seq. analyses | 127% | 125% | 125% |
458.sjeng | Chess | 120% | 151% | 125% |
462.libquantum | Quantum sim | 91% | 129% | 90% |
464.h264ref | Video encoding | 101% | 112% | 112% |
471.omnetpp | Network sim | 109% | 116% | 103% |
473.astar | Pathfinding | 140% | 149% | 137% |
483.xalancbmk | XML processing | 120% | 107% | 116% |
On average, both Xeons pick up about 20% due to SMT (Hyperthreading). The EPYC 7601 improved by even more: it gets a 28% boost on average. There are many possible explanations for this, but two are the most likely. In the situation where AMD's single threaded IPC is very low because it is waiting on the high latency of a further away L3-cache (>8 MB), a second thread makes sure that the CPU resources can be put to better use (like compression, the network sim). Secondly, we saw that AMD core is capable of extracting more memory bandwidth in lightly threaded scenarios. This might help in the benchmarks that stress the DRAM (like video encoding, quantum sim).
Nevertheless, kudos to the AMD engineers. Their first SMT implementation is very well done and offers a tangible throughput increase.
219 Comments
View All Comments
extide - Tuesday, July 11, 2017 - link
PCPer made this same mistake -- Nehalem/Westmere used a crossbar memory bus -- not a ringbus. Only Nehalem/Westmere EX used the ringbus (the 6500/7500 series) The i7 and Xeon 5500 and 5600 series used the crossbar.extide - Tuesday, July 11, 2017 - link
Sandy Bridge brought the ringbus down to Xeon EP and client chips.Yorgos - Tuesday, July 11, 2017 - link
"With the complexity of both server hardware and especially server software, that is very little time. There is still a lot to test and tune, but the general picture is clear."No wonder why we see ubuntu and ancient versions of gcc and the rest of the s/w stack.
Imagine if you tried to use debian or rhel, it would take you decades to get the review.
eligrey - Tuesday, July 11, 2017 - link
Why did you omit the Turbo frequencies for the Xeon Gold 6146 and 6144?Intel ARK says that the 6146's turbo frequency is 4.2GHz and the 6144's is 4.5GHz.
eligrey - Tuesday, July 11, 2017 - link
Oops, I mean 4.2GHz for both.boozed - Tuesday, July 11, 2017 - link
Need more Skylake-SP SKUsrHardware - Tuesday, July 11, 2017 - link
For the purley system, It's listed that you used Chipset Intel Wellsburg B0This information cannot be correct. Lewisburg Chipset is the name of the purley chipset. Also, B0 stepping lewisburg also wouldn't boot with the stepping of CPU you have.
rHardware - Tuesday, July 11, 2017 - link
That 0200011 microcode is also very old.Rickyxds - Tuesday, July 11, 2017 - link
I'am a brazilian processors enthusiast and I'am very critic about intel and AMD processors, between 2012 and Q1 2017 AMD just doesn't existed, who bought AMD on that years, bougth just for love AMD and just it, doesn't for the price, doesn't for the high core count, doesn't for AMD is red, AMD was the worst performance processors. The A9 Apple dual core performance is better than FX 8150.But now I am very surprise with the aggressive AMD prices. No one here Imagined get the Ryzen 7 performance for less than $500. And I don't know if this scenario brings profit to AMD, but for the image against the intel it's wonderful.
On the next years we will see.
krumme - Tuesday, July 11, 2017 - link
Thank you for quality stuff article especially given the short time. So thank you for booting up Johan !Interesting and surpricing results.