GPUs perform exceptionally well at high throughput and low-to-medium interactivity, but their architecture is less suited for ultra-low-latency inference.An 8-GPU HGX B200 server provides a theoretical HBM memory bandwidth of 64 TB/s of in aggrega...