Featured image of post Zen5 and C#

Zen5 and C#

Zen 5 landed with a bit of controversy, but anyone Who was paying any attention to the real results wasn’t disappointed. Benchmarks done using more mature OSes (like Linux and Win10) and professional software shown its strengths. But hey, why should we trust just any 3rd party data if we can make a test by ourselves? I used Ryzen 9950x today and a bit of code to get things going.

Ryzen 9 9950X retail box on a world-map desk mat Ryzen 9 9950X seated in the AM5 socket, retention frame open The assembled board: 9950X under the retention frame, two RAM sticks, Aorus Elite AX heatsinks

Data Processing and Core-to-Core Latency

I am often processing large piles of data. While this is normally executed on different HW than my personal PC, it is always nice to have the computing power locally.

What is counterintuitive in a case of data processing: more CPU cores may not bring better performance. In a CPU heavy tasks (like rendering) this is not apparent, but when the CPU work is combined with synchronization and data transfers the core to core latency, data locality, OS scheduler and other thigs are coming into play.

The Producer–Consumer Test

Setup

I promised a practical example, so lets examine this old problem of producer and consumer. We will dive into details soon (or jump right into the code here), but now lets focus on the main code part. I am using C# with .NET 8.0 and Win10.

    IntPtr mask1 = new(1 << core1);
    IntPtr mask2 = new(1 << core2);

    SemaphoreSlim semaphoreProducer = new(queueSize);
    SemaphoreSlim semaphoreConsumer = new(0);

    string[] message = new string[queueSize];
    int reader = -1;
    int writer = -1;

In the code above we define 2 affinity masks which are binary vectors of ones and zeroes hidden inside single int type. We allow only one core in each mask, values for core1 and core2 are within the range 0..15 (I disabled SMT for this test, otherwise we would have to deal with virtual cores as well). I use two semaphores for tight control of producer and consumer threads. The producer could produce only up to queueSize messages and consumer could consume only as many messages as are available without blocking. Array message is our queue and reader/writer index is used to address next message (or empty spot) in our queue.

Producer

Thread producerThread = new(() =>
{
    ThreadGuard.GetInstance(mask1).Guard();
    for (int i = 1; i <= messagesToProcess; i++)
    {
        semaphoreProducer.Wait();

        int index = Interlocked.Increment(ref writer);
        message[index % queueSize] = i.ToString();

        semaphoreConsumer.Release();
    }
});

Here comes the producer thread. First thing we do is the affinity setting. ThreadGuard implementation is a technical detail for which You have to scroll a bit further, it ensures that the thread it is called from will get executed only on the core(s) we defined by a mask. We produce desired number of messages in a for loop. Each message needs it own place in a queue so we do semaphoreProducer.Wait() as a first thing. Initial capacity of this semaphore is queueSize, so there is plenty of place from the start. We are looping through spots in a queue thanks to modulo operation, thus index % queueSize. The call semaphoreConsumer.Release() is actually telling the consumer about the new message we just prepared.

Consumer

Thread consumerThread = new(() =>
{
    ThreadGuard.GetInstance(mask2).Guard();
    int i = 0;   
    do {
        semaphoreConsumer.Wait();

        int index = Interlocked.Increment(ref reader);
        i = int.Parse(message[index % queueSize]);

        semaphoreProducer.Release();
       } while (i < messagesToProcess);
});

Consumer part looks like a mirror of the producer, we wait in semaphoreConsumer.Wait() until there is some message to consume, once we are done the semaphoreProducer.Release() call signals the producer that the place in queue is now free for a new message. Whole party is going on until we produce and consume desired amount of messages.

ThreadGuard

Complete ThreadGuard code looks like this:

using System;
using System.Collections.Concurrent;
using System.Runtime.InteropServices;

public class ThreadGuard
{
    // Import SetThreadAffinityMask from kernel32.dll
    [DllImport("kernel32.dll")]
    private static extern IntPtr SetThreadAffinityMask(IntPtr hThread, IntPtr dwThreadAffinityMask);

    // Import GetCurrentThread from kernel32.dll
    [DllImport("kernel32.dll")]
    private static extern IntPtr GetCurrentThread();
    
    private static readonly ConcurrentDictionary<IntPtr, ThreadGuard> GuardPool = new();
    
    private readonly IntPtr mask;

    private ThreadGuard(IntPtr mask) { this.mask = mask; }

    public static ThreadGuard GetInstance(IntPtr affinityMask) // factory method
    {
        return GuardPool.GetOrAdd(affinityMask, _ => new ThreadGuard(affinityMask));
    }

    public void Guard()
    {

        IntPtr currentThreadHandle = GetCurrentThread();                  // Get the handle of the current thread
        IntPtr result = SetThreadAffinityMask(currentThreadHandle, mask); // Set the thread affinity mask
        if (result == IntPtr.Zero) { throw new InvalidOperationException("Failed to set thread affinity mask."); }
    }
}

Results

Now the best part - let’s run this for all pairs of cores and various queue sizes. All measured numbers are millions of transferred messages per second. Row numbers on a left are “producer” cores and column numbers on top are “consumer” cores.

Queue Length 1

Table 1 — 250k messages per pair, queue length 1. Millions of messages per second; rows = producer core, columns = consumer core.

CoreA \ CoreB ->0123456789101112131415
00.213.494.053.853.413.743.353.400.820.810.780.800.770.740.840.73
13.890.193.323.763.633.543.644.030.840.900.980.870.880.680.900.70
23.612.520.193.344.083.582.914.010.900.820.931.040.811.050.831.08
34.112.434.040.193.763.543.993.320.890.880.930.880.990.910.990.80
43.083.873.563.550.193.494.023.500.860.950.910.921.090.840.870.88
53.563.483.303.823.480.193.694.110.881.080.990.950.930.900.810.88
63.633.473.453.633.503.460.194.260.940.970.951.090.891.160.950.88
73.053.084.713.263.533.363.490.190.950.910.940.890.930.760.851.01
80.821.160.910.790.870.831.000.790.183.754.003.413.172.993.483.60
90.760.750.820.970.901.040.850.813.890.183.663.633.013.553.173.90
100.801.020.810.820.830.760.770.743.704.450.183.234.773.492.784.11
110.780.870.830.860.890.810.801.003.373.604.040.184.403.583.572.98
120.750.791.010.830.690.820.760.853.253.673.184.230.203.883.783.36
130.890.930.980.841.040.880.900.883.463.543.833.813.390.203.123.81
140.780.830.770.890.840.740.900.783.563.083.393.433.333.470.204.06

All five tables as colour-coded heat maps (green = fastest, yellow = same CCD, orange = cross-CCD), the way they were originally published — the plain tables follow below:

Table 1 as a colour-coded heat map — queue length 1 Table 2 as a colour-coded heat map — queue length 10 Table 3 as a colour-coded heat map — queue length 100 Table 4 as a colour-coded heat map — 25M messages, queue length 10 000 Table 5 as a colour-coded heat map — 25M messages, queue length 1 000

Starting with queue length of 1 we are not letting any thread to produce more than 1 message without blocking. This looks interesting. Even worse than talking to another CCD is the single core trying to execute two things at once. I am surely not alone who hates to do context switching while working on multiple tasks. Solution to this (as visible later) is to at least do more work on a single task before switching to another one. For those who still wonder why is this so bad - we are really not allowing any thread to do any useful work until the other thread gets its fair share of CPU time. Measured throughput of messages is at the same time also measuring number of context switches between consumer and producer threads by the OS when only single CPU core is used. When we shift our attention to the single vs. cross CCD communication using 2 cores, it comes about 5x slower.

Queue Length 10 and 100

Table 2 — 250k messages per pair, queue length 10.

CoreA \ CoreB ->0123456789101112131415
01.635.475.455.255.625.315.374.731.191.231.291.231.291.261.121.35
15.971.765.006.424.575.995.326.181.621.291.241.041.421.001.151.09
25.976.071.755.897.345.825.415.981.351.301.151.441.561.321.461.16
35.724.996.051.765.526.016.255.901.331.291.241.441.231.151.311.29
45.825.566.045.281.765.536.135.921.251.451.321.601.301.641.051.39
55.015.566.096.856.191.735.715.701.421.301.201.361.341.161.321.35
65.615.265.566.005.545.841.745.551.351.641.531.161.161.161.321.36
74.995.625.825.875.515.875.841.731.231.361.451.071.421.391.341.21
81.011.231.271.241.071.341.361.301.695.414.856.105.666.145.355.01
91.251.241.091.251.191.231.271.125.661.706.695.844.965.505.115.04
101.291.231.141.281.391.091.131.085.665.011.695.996.385.907.005.33
110.921.111.301.441.311.101.341.265.395.425.741.705.575.664.865.39
121.291.611.191.091.151.321.001.415.965.955.955.321.695.605.805.36
131.091.041.491.301.471.291.241.244.525.235.376.165.521.695.876.41
141.081.021.131.231.261.191.051.295.964.835.165.385.705.801.665.62
Table 3 — 250k messages per pair, queue length 100.
CoreA \ CoreB ->0123456789101112131415
08.275.545.215.395.145.484.934.821.331.301.341.171.321.361.301.40
16.1910.056.165.316.205.505.724.761.411.351.401.331.341.411.441.41
24.634.5910.006.516.235.436.115.201.581.511.511.621.461.331.251.32
35.755.846.239.925.576.056.025.411.421.181.451.371.121.431.261.15
45.765.916.135.4210.006.435.655.641.571.201.231.331.381.301.151.41
55.336.105.586.276.299.995.315.861.191.311.521.381.372.041.201.33
65.735.305.385.615.936.129.906.001.761.611.181.191.321.511.211.35
75.396.255.695.735.735.985.819.811.291.671.461.461.091.231.321.35
81.331.131.311.431.191.461.321.379.595.996.186.065.885.665.595.56
91.151.341.411.411.361.351.061.216.029.645.415.245.125.055.125.50
101.311.231.001.131.421.331.321.155.476.339.736.005.635.285.775.48
111.171.291.571.251.281.191.601.105.336.295.719.605.836.215.315.37
121.152.111.441.111.201.161.481.466.315.966.185.589.715.615.785.14
131.141.331.561.481.151.751.091.315.935.236.075.696.719.745.665.89
141.121.171.371.351.121.171.301.495.774.605.305.025.645.659.585.84

When allowed to post at least 10 messages into queue at once, single core score surpasses the cross CCD score by a little. With queue length of 100, the single core score is already the fastest.

25 Million Messages

Table 4 — 25M messages per pair, queue length 10 000.

CoreA \ CoreB ->0123456789101112131415
018.725.455.415.105.554.985.034.871.241.331.371.321.321.141.171.35
15.0318.855.225.054.925.144.974.821.131.261.281.361.141.331.101.19
25.515.3719.045.395.225.255.415.271.361.141.351.081.181.191.271.35
35.275.745.7419.145.575.794.935.441.141.231.121.321.361.161.331.26
45.375.405.395.3619.005.815.575.341.181.141.021.341.141.111.331.04
55.385.375.405.565.5019.015.075.621.321.341.031.321.161.341.101.31
65.345.105.155.425.665.5618.775.681.301.341.351.181.131.351.261.02
75.165.185.075.475.325.355.3918.831.111.021.171.101.341.041.251.34
81.041.061.340.991.021.131.351.0218.555.645.585.345.505.325.014.83
91.021.171.350.981.321.171.020.995.5718.525.355.644.915.074.714.75
101.001.111.201.301.361.331.331.175.025.2618.475.515.585.295.354.94
111.361.341.241.001.341.121.301.375.295.435.4618.565.415.274.965.18
121.331.141.231.121.191.161.311.195.275.265.455.2918.485.645.335.46
131.311.131.341.101.191.161.081.115.075.235.235.555.7018.505.345.40
141.341.321.101.101.331.331.101.325.034.885.045.105.295.2918.275.21
Table 5 — 25M messages per pair, queue length 1 000 (54 min run).
CoreA \ CoreB ->0123456789101112131415
017.245.295.585.325.494.895.194.591.311.311.191.261.421.071.211.24
15.3917.445.335.945.305.624.815.031.191.141.171.161.271.181.171.12
25.095.2417.366.015.315.225.165.211.191.211.031.121.331.331.161.31
35.255.615.8117.745.155.024.985.190.951.231.091.191.121.161.171.03
45.265.395.585.4017.315.395.255.121.341.211.091.161.031.191.301.29
55.205.285.295.515.4817.315.365.491.131.291.341.351.131.311.171.14
64.995.224.995.015.415.2717.525.361.281.321.041.141.301.311.071.31
75.165.194.845.555.175.535.4617.181.271.270.981.311.351.161.181.11
81.331.111.201.041.271.151.031.3316.785.425.245.415.235.254.794.92
91.301.311.021.291.351.351.081.135.3516.915.195.235.005.404.695.28
101.180.941.051.301.311.371.161.025.264.9617.375.315.675.325.064.76
111.121.121.291.141.341.301.091.155.055.045.4316.945.345.084.804.96
121.171.211.201.251.171.121.291.265.135.115.495.0616.875.205.204.86
131.051.331.151.361.151.141.041.224.915.155.485.335.2916.684.875.07
141.101.341.211.321.121.331.051.084.855.124.824.655.025.0316.615.02

Finally the 25M messages exchanged for each pair with queue length of 1k and 10k messages. It took almost an hour to collect each. Single core, single CCD and cross CCD latencies are clearly separated by our own test.

Windows Task Manager during a test run on the Ryzen 9 9950X: two of the sixteen cores fully loaded, the rest idle

Conclusion

It didn’t took any sophisticated equipment or software, yet We were able to gain valuable insight into performance characteristics of a modern multi core and multi CCD CPU.

All code used for this article is available on GitHub.

If You have any questions, ideas or if You believe I am completely wrong, please leave me a comment on dev.to!

References

  1. ccd2ccd-messaging, the producer–consumer benchmark used in this post (C#), GitHub.
Questions, ideas, or think I got it wrong? This article is also published on dev.to — leave a comment there.
Built with Hugo
Theme Stack designed by Jimmy