Prefer to read without ads? Become a member — from $10/month — and support the work. Already a member? Log in to read ad-free on this device.

Further Reading

NVIDIA has white papers on their Web site that describe the Fermi and Kepler architectures in detail. This white paper describes Fermi:

The next generation of NVIDIA GeForce GPU. http://www.nvidia.com/object/GTX_400_architecture.html

This white paper describes the Kepler architecture and its implementation in the NVIDIA GeForce GTX 680 (GK104):

https://www.nvidia.com/content/PDF/product-specifications/GeForce_GTX_680_Whitepaper_FINAL.pdf

NVIDIA has published a comparable architecture white paper for every subsequent generation—Maxwell, Pascal, Volta, Turing, Ampere, Ada Lovelace, Hopper, and Blackwell—each available from NVIDIA’s Web site and each the most detailed public description of its generation’s Streaming Multiprocessors and memory subsystem.

NVIDIA engineers also have published several architectural papers that give more-detailed descriptions of the various CUDA-capable GPUs:

Lindholm, E., J. Nickolls, S. Oberman, and J. Montrym. NVIDIA Tesla: A unified graphics and computing architecture. IEEE Micro 28 (2), March-April 2008, pp. 39-55.

Wittenbrink, C., E. Kilgariff and A. Prabhu. Fermi GF100 GPU Architecture. IEEE Micro 31 (2), March-April 2011, pp. 50-59.

Foley, D., and J. Danskin. Ultra-Performance Pascal GPU and NVLink Interconnect. IEEE Micro 37 (2), March-April 2017, pp. 7-17.

Choquette, J., O. Giroux, and D. Foley. Volta: Performance and Programmability. IEEE Micro 38 (2), March-April 2018, pp. 42-52.

Choquette, J., W. Gandhi, O. Giroux, N. Stam, and R. Krashinsky. NVIDIA A100 Tensor Core GPU: Performance and Innovation. IEEE Micro 41 (2), March-April 2021, pp. 29-35.

Choquette, J. NVIDIA Hopper H100 GPU: Scaling Performance. IEEE Micro 43 (3), May-June 2023, pp. 9-17.

Wong et al. used CUDA to develop microbenchmarks and clarify some aspects of Tesla-class hardware architecture:

Wong, H., M. Papadopoulou, M. Sadooghi-Alvandi, and A. Moshovos. Demystifying GPU Microarchitecture through microbenchmarking. 2010 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 28-30 March 2010, pp. 235-246.

Researchers at Citadel have continued this line of work, using microbenchmarks to reverse-engineer the instruction sets, memory hierarchies, and Tensor Cores of more recent architectures:

Jia, Z., M. Maggioni, B. Staiger, and D.P. Scarpazza. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking. arXiv:1804.06826, 2018. https://arxiv.org/abs/1804.06826

Jia, Z., M. Maggioni, J. Smith, and D.P. Scarpazza. Dissecting the NVIDIA Turing T4 GPU via Microbenchmarking. arXiv:1903.07486, 2019. https://arxiv.org/abs/1903.07486

Other groups have applied the same microbenchmarking approach to the GPU memory hierarchy and to a succession of architectures from Maxwell onward:

Zhang, X., G. Tan, S. Xue, J. Li, K. Zhou, and M. Chen. Understanding the GPU Microarchitecture to Achieve Bare-Metal Performance Tuning. Proceedings of the 22nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP), 2017, pp. 31-43.

Mei, X., and X. Chu. Dissecting GPU Memory Hierarchy through Microbenchmarking. IEEE Transactions on Parallel and Distributed Systems 28 (1), January 2017, pp. 72-86.

Abdelkhalik, H., Y. Arafa, N. Santhi, and A.-H. Badawy. Demystifying the Nvidia Ampere Architecture through Microbenchmarking and Instruction-level Analysis. 2022 IEEE High Performance Extreme Computing Conference (HPEC), September 2022, pp. 1-8. https://arxiv.org/abs/2208.11174

Luo, W., R. Fan, Z. Li, D. Du, H. Liu, Q. Wang, and X. Chu. Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis. arXiv:2501.12084, 2025. https://arxiv.org/abs/2501.12084

Jarmusch, A., N. Graddon, and S. Chandrasekaran. Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks. arXiv:2507.10789, 2025. https://arxiv.org/abs/2507.10789

For the interconnects covered in Section 2.7, Li et al. benchmarked PCIe, NVLink, NV-SLI, NVSwitch, and GPUDirect head-to-head across single- and multi-GPU topologies:

Li, A., S. Song, J. Chen, J. Li, X. Liu, N.R. Tallent, and K.J. Barker. Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect. IEEE Transactions on Parallel and Distributed Systems 31 (1), January 2020, pp. 94-110. https://arxiv.org/abs/1903.04611