Computer Music - Hardware,Test Labs DAWBench Threadripper Pro Testing – 9965WX & 9975WX

DAWBench Threadripper Pro Testing – 9965WX & 9975WX

AMD Threadripper Pro

In the world of high-end recording solutions, whilst workstation grade CPU offerings have often been in the running, the past few generations have failed to set the world on fire in terms of studio practicality, especially given that the more value orientated consumer chips have continued to make sizable performance gains over this timeframe. It hasn’t just been the cost to performance ratio that’s in question, rather the various aspects of these platforms such as a focus on core count over single core performance, which have resulted in solutions that in general haven’t always lined-up with our audio workload requirements. This has helped to ensure that at least in audio terms, the spotlight has remained more focused upon the consumer level components for the past couple of years.

AMD came closest to breaking this pattern with their last 7000 series refresh, although overall the solution still fell short on some aspects. This time around we see another silicon overhaul as AMDs Zen 5 arrives on the Threadripper platform, bringing with it improvements that may well help to resolve those earlier sticking points.

The previous generation experienced a performance bottleneck similar to problems seen before in the early Ryzen consumer implementations, which would result in our ASIO 64 buffer (and lower) struggling during more memory intensive VSTi library testing. With the testing suggesting that there was an element of internal latency, the platform found itself better suited to off-line render processing or other less latency dependent real-time workloads. The 9000 series sees the Infinity Fabric Bus get a speed bump from 6000MHz to 6400MHz this time around and notably, with similar memory speed support also having been added to this new release. Along with this we’ve seen ever faster memory kits becoming available on the wider market, this should help us in theory close the low-latency performance gap. The question as we entered into testing was very much one of has it done enough to turn this into the ultimate studio platform?

In terms of CPUs we’re looking at the 9965WX (24 core / 48 Thread) and 9975WX (32 core / 64 thread) at this time. The 9975WX spec includes a base clock of 4GHz and up to 5.4GHz turbo depending on workload, whereas the 9985WX above it reduces this to 3.2GHz base clock, less than ideal when we’re looking to strike a balance between single core and multi-core performance.

The mainboard in use is the ASUS SAGE WX90E with some additional highlighted memory testing on a ASUS SAGE TRX50, both cooled using a Thermaltake Thermaltake AW420 420mm AIO cooler and powered by a 1600W Seasonic Prime TX.

DAWBench DSP Threadripper Chart

The DSP testing in particular has previously come away with favourable results for the Threadripper chips and this time around sees no real surprise, as both of the chips here both pull away strongly from the pack. We see around a 40% – 45% jump in performance when considering the 24 core 9965WX CPU, whereas the 9975WX results in this test look to be somewhere in the 80% – 90% gains performance wise over the Intel 285K consumer level CPU leader. Performance here continues to scale efficiently across the ASIO buffer settings and the DSP reflects performance for in the box style generative synths, such as Phase Plant, Pigments & Serum or any of the general VST effects processing applied within your projects.

DAWBench VI Threadripper Chart

In contrast, the low-latency weakness in the platform would typically show up in the VI testing when run at the tightest of ASIO buffers and we continue to see a lull in the results here, although it’s proved to be an improvement over the prior generation testing. The VI Kontakt based test focuses on sample based playback which relies on both the general memory performance, along with the CPU performance and the general interaction between the two. We’ve seen this occur on older generations previously and whilst it’s a issue that’s long since been cleared up within the consumer range, the trailing Threadripper hardware has AMD still working to iron this out.

Where it may prove more difficult for the Threadripper chips is the increase in CCD’s, AMD’s collection of cores in a block layout design. When working with multi-CCD chips we can expect to see a degree of latency occurring between the CCD blocks, although even the largest of the consumer grade chips the 9950X has a limit of 2 CCD’s to contend with, whereas the CPUs we see here are based around a quad CCD design, with the chips above these at the top of the Threadripper range expanding this to 8 CCD blocks in use.

Whilst the OS will often attempt to store data as locally as possible to a given core, as projects grow and data is shared around, it’s inevitable that data required by CCD block A may need to be called from a memory space attached to CCD block D, leading to a certain amount of latency coming into play as that data is recalled. The internal bus speed is crucial here and memory speed can also help to mitigate this to some extent, where optimizing it to run 1:1 against the internal FCLK setting is a common optimization advised with AMD based configurations. When looking at the prior generation, it was early in the initial release cycle and high-speed EEC based DDR5 memory kits were still a long way from reaching the general market, in addition officially the 7000 series chips were rated to support 5200MHz with a 6000MHz Infinity Fabric. BIOS level support increased over the generational cycle opening up the options for faster RAM kits as the individual board manufacturers validated faster kits over time, however it would still ultimately be an overclock compared with CPUs memory controller rating.

This time around a growing number of AMD EXPO supporting kits have reached the market and we are slowly starting to see larger, faster kits appearing more in line with the supported rating of 6400MHz offered by the latest 9000 series chips. Even so, in quantities large enough to support the VI test here, we were limited to 6000MHz at the time of testing, but even then this looks to have helped close the latency gap previously seen. Whereas the 7000 series would often fail entirely when running the RME Babyface in testing at it’s 64 ASIO buffer setting, this time around we see an outcome returned that whilst low given the overall performance on offer by the chip, it is now returning a result that is nestled amongst the lower end of the chart.

However, once we move beyond this to just the 128 buffer setting, we see the performance begin to realign as to where you would more naturally expect with the 9650WX seeing roughly 80% increase over the next nearest CPU and the 9750WX offering some 20% – 30% extra beyond that, easily doubling the performance the Intel Ultra 285K.

Not ideal for everyone as the VI test showing up the performance hole with memory intensive workloads and real-time processing may prove problematic for some users, this may include live performance examples where you would be looking to process live audio on the fly or trigger and apply audio from your sample libraries in a live environment. These tend to be situations where not only is the RAM performance important, but then the CPUs single core performance is also crucial and it might be argued that a solution built around a far cheaper, but more single core performance optimized solution like an Intel Ultra series chip or one of AMDs own Ryzen solutions would likely offer a far more suitable solution for such a job. However, for those working largely in the box, doing mixing, mastering or other post-production jobs then this latency tends to be less important, or at least you will be able to run it at a more relaxed level which in turn would allow you to take full advantage of the overhead available here.

Having noted the remaining lull in performance still appearing at certain settings, I took the opportunity to run up some additional RAM testing where we saw some interesting patterns appear.

Dawbench TR Memory Chart

The uppermost result is using the 9975WX chip, but running it instead on the TRX board designed for the HDET X series chips. The Threadripper chip design in contrast to the consumer level hardware, will offer up dedicated channel per RAM stick and whilst both the TRX and WRX boards may support the same WX range chips, the TRX boards will only offer 4 slots with 4 channels, whereas the WRX boards offer up to 8 slots and 8 channels of RAM support. If we run a WX level CPU in a TRX board, then it only receives half of its possible bandwidth, which looks to show up to some degree with the VI testing. In contrast the two results on the WRX board at 5600MHz and 6000MHz are both fully populated with 128GB’s worth of memory across 8 sticks and all kits running at AMD EXPO settings.

The results highlight the benefits about maxing out the RAM bandwidth in this test through running the full 8 channel configuration where possible. We also get to see where the bump in memory speed from 5600MHz to 6000MHz further helping to improve the polyphony results. The sweet spot with this generation of chips is noted to be at 6400MHz and whilst I didn’t have the suitable kits to try this out across all 8 channels, we typically have seen the optimal performance with AMD CPUs when matching the Infinity Fabric at 1:1 ratio. As more AMD EXPO ready kits continue to reach the market and we see more choice, looking to fit memory around this speed rating is likely to help you get the best possible results out of the platform.

Lastly we arrive at the DAWBench BUS test, with its focus on single core performance and to some degree the internal signal routing.

DAWBench BUS Threadripper Chart

This ends up being another factor which may prove less than desirable for real-time processing situations as the internal latency looks to rear its head again here, but the results do begin to recover again on the 128 ASIO buffer setting and they quickly draw level again with the rest of the Zen 5 based chips on the chart once we reach the 256 buffer and above.

The release of the Threadripper 9000 series in audio terms looks to be a marked improvement over the prior generation and a platform that continues to make strides towards reaching its full potential, although it’s still not quite the ultimate all-rounder. Just as with the consumer desktop chips before them, up until the AMD 5000 desktop series we saw incremental gains with RAM performance through each iteration along with the tightening up of the internal CPU bus clocking. With this platform also slowly closing the gap on each new release, further improvements to forthcoming silicon and ever increasing RAM performance suggests this is a platform that’s getting close to reaching its full potential in terms of audio workflows. As it is, it does have a few weaknesses shown up in testing, although ones that honestly might not be of a concern to a large number of potential users of the platform. The fact is, where it does excel, it does so to a highly impressive degree, certainly where it would prove to be an outstanding solution in many workflows given the right considerations.

Scan 3XS Custom Audio Systems
Scan Fixed Series Audio Systems