Prefer to read without ads? Become a member — from $10/month — and support the work. Already a member? Log in to read ad-free on this device.

11.5 Performance and Summary

This chapter covered four different implementations of SAXPY, emphasizing different strategies of data movement:

Table 11-1 summarizes the relative performance of these implementations for 128M floats on a GeForce RTX 3060. Each successive strategy improves on the last, and—despite issuing no asynchronous memcpy calls of its own—mapped pinned memory delivers the highest effective bandwidth.

Version Bandwidth (MB/s)
stream1Device.cu 18343
stream2Async.cu 24652
stream3Streams.cu 31537
stream4Mapped.cu 40074

Table 11-1. SAXPY streaming performance (128M floats, GeForce RTX 3060)