Hi,
as previously described here, out-of-box achievable latency for the RoCE loopback test using HSB 2.7.0 with a GB10-powered machine was 5.57μs.
This is pretty close to the 4μs latency required for QEC; close enough that it is worth exploring whether optimizations to the data flow might be worth pursuing.
First question would be, is it reasonable to expect that such a latency (4μs max) could be achieved, in the considered use case, assuming the minimal possible (software) overhead?
Second question would be, would you have any suggestions on what to pursue, and, more importantly, in which order?
The current thinking on our side is that, since GB10 cannot do GPUdirect, the lowest hanging fruit would be to move all things related to buffers management into the kernel (for proof-of-concept, this'd mean, direct changes to the drivers; if it works, and if it would be feasible to do so, extract the code into a standalone kernel module).
Assuming that made a difference, another venue we'd consider is, running a "minimal-noise" environment - potentially reducing the in-use ARM core count, while using the smallest possible userland - the general idea being, to minimize the pressure put on RAM by the cores, by keeping as much as possible in the cache hierarchy.
Any comments, suggestions or advice would be appreciated.
Best regards,
MI
Hi,
as previously described here, out-of-box achievable latency for the RoCE loopback test using HSB 2.7.0 with a GB10-powered machine was 5.57μs.
This is pretty close to the 4μs latency required for QEC; close enough that it is worth exploring whether optimizations to the data flow might be worth pursuing.
First question would be, is it reasonable to expect that such a latency (4μs max) could be achieved, in the considered use case, assuming the minimal possible (software) overhead?
Second question would be, would you have any suggestions on what to pursue, and, more importantly, in which order?
The current thinking on our side is that, since GB10 cannot do GPUdirect, the lowest hanging fruit would be to move all things related to buffers management into the kernel (for proof-of-concept, this'd mean, direct changes to the drivers; if it works, and if it would be feasible to do so, extract the code into a standalone kernel module).
Assuming that made a difference, another venue we'd consider is, running a "minimal-noise" environment - potentially reducing the in-use ARM core count, while using the smallest possible userland - the general idea being, to minimize the pressure put on RAM by the cores, by keeping as much as possible in the cache hierarchy.
Any comments, suggestions or advice would be appreciated.
Best regards,
MI