This chapter covered four different implementations of SAXPY, emphasizing different strategies of data movement:
Synchronous memcpy to and from device memory,
Asynchronous memcpy to and from device memory,
Asynchronous memcpy using streams, and
Mapped pinned memory.
Table 11-1 summarizes the relative performance of these implementations for 128M floats on a GeForce RTX 3060. Each successive strategy improves on the last, and—despite issuing no asynchronous memcpy calls of its own—mapped pinned memory delivers the highest effective bandwidth.
| Version | Bandwidth (MB/s) |
|---|---|
stream1Device.cu |
18343 |
stream2Async.cu |
24652 |
stream3Streams.cu |
31537 |
stream4Mapped.cu |
40074 |
Table 11-1. SAXPY streaming performance (128M floats, GeForce RTX 3060)