1 result found
What changed
CUDA reconstruction pipeline was memory-bandwidth bound, with sub-linear multi-GPU scaling eating into throughput.
What we delivered
Profiling-driven optimizations plus a CI-backed performance/correctness harness, revisited and re-tuned as the client moved to newer GPU generations.