Prefer to read without ads? Become a member — from $10/month — and support the work. Already a member? Log in to read ad-free on this device.

8.7 Instruction Sets

Since the first CUDA-capable hardware shipped in 2006, NVIDIA has developed eleven major architectures – Tesla, Fermi, Kepler, Maxwell, Pascal, Volta, Turing, Ampere, Ada Lovelace, Hopper, and Blackwell (Table 8-1) – and the native instruction set has grown with each of them, from a few dozen operations to well over a hundred per generation. New instructions appear both between architectures and, occasionally, within a family as NVIDIA refreshes its products. Global atomic operations, for example, were absent from the very first Tesla-class processor (the G80, which shipped in 2006 as the GeForce 8800 GTX) but present in every Tesla GPU thereafter: querying the compute capability returns 1.0 for the G80 and 1.1 or greater for the rest, so an application can use global atomics whenever the reported version is at least 1.1.

The native instructions are called SASS. CUDA does not document SASS in the way it documents PTX: NVIDIA publishes no encodings or semantics for it, reserves the right to change it from one architecture to the next, and offers no guarantee of binary compatibility across generations – which is precisely what lets the instruction set evolve as freely as it has. What NVIDIA does publish, in the CUDA Binary Utilities documentation, is a per-architecture list of the instruction mnemonics that cuobjdump and nvdisasm emit when disassembling a binary, each with a one-line description. That reference is the authoritative, living catalog of the SASS instruction set for every architecture from Turing (SM 7.5) forward. The tables that follow give the SASS instruction set of every architecture from Tesla through Blackwell – from that reference for Turing onward, from the edition archived with CUDA 10.0 for Fermi through Volta, and from the first edition of this book for the Tesla generation, which predates it.

The instruction sets below give the SASS mnemonics for every CUDA architecture, grouped by instruction category, with the opcode as it appears in disassembly. Fermi through Volta are drawn from the CUDA Binary Utilities reference archived with CUDA 10.0; Turing through Blackwell from the current edition; and the Tesla set, which predates both, from the first edition of this book, whose middle column gives the first SM version to support each instruction.

Tesla (SM 1.0 / 1.1 / 1.3)

Floating Point

Opcode SM Description
COS 1.0 Cosine
DADD 1.3 Double-precision floating point add
DFMA 1.3 Double-precision floating point fused multiply-add
DMAX 1.3 Double-precision floating point maximum
DMIN 1.3 Double-precision floating point minimum
DMUL 1.3 Double-precision floating point multiply
DSET 1.3 Double-precision floating point condition set
EX2 1.0 Exponential (base 2)
FADD/FADD32/FADD32I 1.0 Single-precision floating point add
FCMP 1.0 Single-precision floating point compare
FMAD/FMAD32/FMAD32I 1.0 Single-precision floating point multiply-add*
FMAX 1.0 Single-precision floating point maximum
FMIN 1.0 Single-precision floating point minimum
FMUL/FMUL32/FMUL32I 1.0 Single-precision floating point multiply
FSET 1.0 Single-precision floating point conditional set
LG2 1.0 Single-precision floating point logarithm (base 2)
RCP 1.0 Single-precision floating point reciprocal
RRO 1.0 Range reduction operator (used before SIN/COS)
RSQ 1.0 Reciprocal square root
SIN 1.0 Sine

Flow Control

Opcode SM Description
BAR 1.0 Barrier synchronization/ __syncthreads()
BRA 1.0 Conditional branch
BRK 1.0 Conditional break from loop
BRX 1.0 Fetch an address from constant memory and branch to it
C2R 1.0 Condition code to data register
CAL 1.0 Unconditional subroutine call
RET 1.0 Conditional return from subroutine
SSY 1.0 Set synchronization point; used before potentially divergent instructions

Data Conversion

Opcode SM Description
F2F 1.0 Copy floating point value with conversion to floating point
F2I 1.0 Copy floating point value with conversion to integer
I2F 1.0 Copy integer value to floating-point with conversion
I2I 1.0 Copy integer value to integer with conversion

Integer

Opcode SM Description
IADD/ IADD32/ IADD32I 1.0 Integer addition
IMAD/ IMAD32/ IMAD32I 1.0 Integer multiply-add
IMAX 1.0 Integer maximum
IMIN 1.0 Integer minimum
IMUL/ IMUL32/ IMUL32I 1.0 Integer multiply
ISAD/ ISAD32 1.0 Integer sum of absolute difference
ISET 1.0 Integer conditional set
SHL 1.0 Shift left
SHR 1.0 Shift right

Memory Operations

Opcode SM Description
A2R 1.0 Move address register to data register
ADA 1.0 Add immediate to address register
G2R 1.0 Move from shared memory to register. A .LCK suffix indicates that the bank is locked until an R2G.UNL has been performed; this is used to implement shared memory atomics.
GATOM.IADD/ EXCH/ CAS/ IMIN/ IMAX/ INC/ DEC/ IAND/ IOR/ IXOR 1.2 Global memory atomic operations; performs an atomic operation and returns the original value.
GLD 1.0 Load from global memory
GRED.IADD/ IMIN/ IMAX/ INC/ DEC/ IAND/ IOR/ IXOR 1.2 Global memory reduction operations; performs an atomic operation with no return value.
GST 1.0 Store to global memory
LLD 1.0 Load from local memory
LST 1.0 Store to local memory
LOP 1.0 Logical operation (AND/OR/XOR)
MOV/ MOV32 1.0 Move source to destination
MVC 1.0 Move from constant memory
MVI 1.0 Move immediate
R2A 1.0 Move register to address register
R2C 1.0 Move data register to condition code
R2G 1.0 Store to shared memory. When used with the .UNL suffix, releases a previously held lock on that shared memory bank.

Miscellaneous

Opcode SM Description
NOP 1.0 No operation
TEX/ TEX32 1.0 Texture fetch
VOTE 1.2 Warp-vote primitive.
S2R 1.0 Move special register (e.g., thread ID) to register
S2R 1.0 Move special register (e.g., thread ID) to register

Fermi (SM 2.0 / 2.1)

Floating Point Instructions

Opcode Description
FFMA FP32 Fused Multiply Add
FADD FP32 Add
FCMP FP32 Compare
FMUL FP32 Multiply
FMNMX FP32 Minimum/Maximum
FSWZ FP32 Swizzle
FSET FP32 Set
FSETP FP32 Set Predicate
FCHK FP32 Division Test
RRO FP Range Reduction Operator
MUFU FP Multi-Function Operator
DFMA FP64 Fused Multiply Add
DADD FP64 Add
DMUL FP64 Multiply
DMNMX FP64 Minimum/Maximum
DSET FP64 Set
DSETP FP64 Set Predicate

Integer Instructions

Opcode Description
IMAD Integer Multiply Add
IMADSP Integer Extract Multiply Add
IMUL Integer Multiply
IADD Integer Add
ISCADD Integer Scaled Add
ISAD Integer Sum Of Abs Diff
IMNMX Integer Minimum/Maximum
BFE Integer Bit Field Extract
BFI Integer Bit Field Insert
SHR Integer Shift Right
SHL Integer Shift Left
SHF Integer Funnel Shift
LOP Integer Logic Op
FLO Integer Find Leading One
ISET Integer Set
ISETP Integer Set Predicate
ICMP Integer Compare and Select
POPC Population Count

Conversion Instructions

Opcode Description
F2F Float to Float
F2I Float to Integer
I2F Integer to Float
I2I Integer to Integer

Movement Instructions

Opcode Description
MOV Move
SEL Conditional Select/Move
PRMT Permute
SHFL Warp Shuffle

Predicate/CC Instructions

Opcode Description
P2R Predicate to Register
R2P Register to Predicate
CSET CC Set
CSETP CC Set Predicate
PSET Predicate Set
PSETP Predicate Set Predicate

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4 Texels
TXQ Texture Query

Compute Load/Store Instructions

Opcode Description
LDC Load from Constant
LD Load from Memory
LDG Non-coherent Global Memory Load
LDL Load from Local Memory
LDS Load from Shared Memory
LDSLK Load from Shared Memory and Lock
ST Store to Memory
STL Store to Local Memory
STS Store to Shared Memory
STSCUL Store to Shared Memory Conditionally and Unlock
ATOM Atomic Memory Operation
RED Atomic Memory Reduction Operation
CCTL Cache Control
CCTLL Cache Control (Local)
MEMBAR Memory Barrier

Surface Memory Instructions

Opcode Description
SUCLAMP Surface Clamp
SUBFM Surface Bit Field Merge
SUEAU Surface Effective Address
SULDGA Surface Load Generic Address
SUSTGA Surface Store Generic Address

Control Instructions

Opcode Description
BRA Branch to Relative Address
BRX Branch to Relative Indexed Address
JMP Jump to Absolute Address
JMX Jump to Absolute Indexed Address
CAL Call to Relative Address
JCAL Call to Absolute Address
RET Return from Call
BRK Break from Loop
CONT Continue in Loop
SSY Set Sync Relative Address
PBK Pre-Break Relative Address
PCNT Pre-Continue Relative Address
PRET Pre-Return Relative Address
BPT Breakpoint/Trap
EXIT Exit Program

Miscellaneous Instructions

Opcode Description
NOP No Operation
S2R Special Register to Register
B2R Barrier to Register
BAR Barrier Synchronization
VOTE Query condition across threads

Kepler (SM 3.0 / 3.5 / 3.7)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FCHK Single Precision FP Divide Range Check
FCMP FP32 Compare to Zero and Select Source
FFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMUL FP32 Multiply
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Add used for FSWZ emulation
MUFU Multi Function Operation
RRO Range Reduction Operator FP
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMNMX FP64 Minimum/Maximum
DMUL FP64 Multiply
DSET FP64 Compare And Set
DSETP FP64 Compare And Set Predicate
HADD2 FP16 Add
HFMA2 FP16 Fused Mutiply Add
HMUL2 FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate

Integer Instructions

Opcode Description
BFE Bit Field Extract
BFI Bit Field Insert
FLO Find Leading One
IADD Integer Addition
IADD3 3-input Integer Addition
ICMP Integer Compare to Zero and Select Source
IMAD Integer Multiply And Add
IMADSP Extracted Integer Multiply And Add.
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
ISCADD Scaled Integer Addition
ISET Integer Compare And Set
ISETP Integer Compare And Set Predicate
LEA Compute Effective Address
LOP Logic Operation
LOP3 3-input Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
XMAD Integer Short Multiply Add

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion

Movement Instructions

Opcode Description
MOV Move
PRMT Permute Register Pair
SEL Select Source with Predicate
SHFL Warp Wide Register Shuffle

Predicate/CC Instructions

Opcode Description
CSET Test Condition Code And Set
CSETP Test Condition Code and Set Predicate
PSET Combine Predicates and Set
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate/CC Register

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TXQ Texture Query
TEXS Texture Fetch with scalar/non-vec4 source/destinations
TLD4S Texture Load 4 with scalar/non-vec4 source/destinations
TLDS Texture Load with scalar/non-vec4 source/destinations

Compute Load/Store Instructions

Opcode Description
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDL Load within Local Memory Window
LDS Local within Shared Memory Window
ST Store to generic Memory
STG Store to global Memory
STL Store within Local or Shared Window
STS Store within Local or Shared Window
ATOM Atomic Operation on generic Memory
ATOMS Atomic Operation on Shared Memory
RED Reduction Operation on generic Memory
CCTL Cache Control
CCTLL Cache Control
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Surface Memory Instructions

Opcode Description
SUATOM Surface Reduction
SULD Surface Load
SURED Atomic Reduction on surface memory
SUST Surface Store

Control Instructions

Opcode Description
BRA Relative Branch
BRX Relative Branch Indirect
JMP Absolute Jump
JMX Absolute Jump Indirect
SSY Set Synchronization Point
SYNC Converge threads after conditional branch
CAL Relative Call
JCAL Absolute Call
PRET Pre-Return From Subroutine
RET Return From Subroutine
BRK Break
PBK Pre-Break
CONT Continue
PCNT Pre-continue
EXIT Exit Program
PEXIT Pre-Exit
BPT BreakPoint/Trap

Miscellaneous Instructions

Opcode Description
NOP No Operation
CS2R Move Special Register to Register
S2R Move Special Register to Register
B2R Move Barrier To Register
BAR Barrier Synchronization
R2B Move Register to Barrier
VOTE Vote Across SIMD Thread Group

Maxwell and Pascal (SM 5.x / 6.x)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FADD32I FP32 Add
FCHK Floating-point Range Check
FFMA32I FP32 Fused Multiply and Add
FFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMUL FP32 Multiply
FMUL32I FP32 Multiply
FSEL Floating Point Select
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Swizzle Add
MUFU FP32 Multi Function Operation
HADD2 FP16 Add
HADD2_32I FP16 Add
HFMA2 FP16 Fused Mutiply Add
HFMA2_32I FP16 Fused Mutiply Add
HMMA Matrix Multiply and Accumulate
HMUL2 FP16 Multiply
HMUL2_32I FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMUL FP64 Multiply
DSETP FP64 Compare And Set Predicate

Integer Instructions

Opcode Description
BMSK Bitfield Mask
BREV Bit Reverse
FLO Find Leading One
IABS Integer Absolute Value
IADD Integer Addition
IADD3 3-input Integer Addition
IADD32I Integer Addition
IDP Integer Dot Product and Accumulate
IDP4A Integer Dot Product and Accumulate
IMAD Integer Multiply And Add
IMMA Integer Matrix Multiply and Accumulate
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
IMUL32I Integer Multiply
ISCADD Scaled Integer Addition
ISCADD32I Scaled Integer Addition
ISETP Integer Compare And Set Predicate
LEA LOAD Effective Address
LOP Logic Operation
LOP3 Logic Operation
LOP32I Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
VABSDIFF Absolute Difference
VABSDIFF4 Absolute Difference

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion
I2IP Integer To Integer Conversion and Packing
FRND Round To Integer

Movement Instructions

Opcode Description
MOV Move
MOV32I Move
PRMT Permute Register Pair
SEL Select Source with Predicate
SGXT Sign Extend
SHFL Warp Wide Register Shuffle

Predicate Instructions

Opcode Description
PLOP3 Predicate Logic Operation
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate Register

Load/Store Instructions

Opcode Description
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDL Load within Local Memory Window
LDS Load within Shared Memory Window
ST Store to Generic Memory
STG Store to Global Memory
STL Store within Local or Shared Window
STS Store within Local or Shared Window
MATCH Match Register Values Across Thread Group
QSPC Query Space
ATOM Atomic Operation on Generic Memory
ATOMS Atomic Operation on Shared Memory
ATOMG Atomic Operation on Global Memory
RED Reduction Operation on Generic Memory
CCTL Cache Control
CCTLL Cache Control
ERRBAR Error Barrier
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TMML Texture MipMap Level
TXD Texture Fetch With Derivatives
TXQ Texture Query

Surface Instructions

Opcode Description
SUATOM Surface Reduction
SULD Surface Load
SURED Atomic Reduction on surface memory
SUST Surface Store

Control Instructions

Opcode Description
BMOV Move Convergence Barrier State
BPT BreakPoint/Trap
BRA Relative Branch
BREAK Break out of the Specified Convergence Barrier
BRX Relative Branch Indirect
BSSY Barrier Set Convergence Synchronization Point
BSYNC Synchronize Threads on a Convergence Barrier
CALL Call Function
EXIT Exit Program
JMP Absolute Jump
JMX Absolute Jump Indirect
KILL Kill Thread
NANOSLEEP Suspend Execution
RET Return From Subroutine
RPCMOV PC Register Move
RTT Return From Trap
WARPSYNC Synchronize Threads in Warp
YIELD Yield Control

Miscellaneous Instructions

Opcode Description
B2R Move Barrier To Register
BAR Barrier Synchronization
CS2R Move Special Register to Register
CSMTEST Clip State Machine Test and Update
DEPBAR Dependency Barrier
GETLMEMBASE Get Local Memory Base Address
LEPC Load Effective PC
NOP No Operation
PMTRIG Performance Monitor Trigger
R2B Move Register to Barrier
S2R Move Special Register to Register
SETCTAID Set CTA ID
SETLMEMBASE Set Local Memory Base Address
VOTE Vote Across SIMD Thread Group
VOTE_VTG Clip State Machine Test and Update

Volta (SM 7.0)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FADD32I FP32 Add
FCHK Floating-point Range Check
FFMA32I FP32 Fused Multiply and Add
FFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMUL FP32 Multiply
FMUL32I FP32 Multiply
FSEL Floating Point Select
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Swizzle Add
MUFU FP32 Multi Function Operation
HADD2 FP16 Add
HADD2_32I FP16 Add
HFMA2 FP16 Fused Mutiply Add
HFMA2_32I FP16 Fused Mutiply Add
HMMA Matrix Multiply and Accumulate
HMUL2 FP16 Multiply
HMUL2_32I FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMUL FP64 Multiply
DSETP FP64 Compare And Set Predicate

Integer Instructions

Opcode Description
BMMA Bit Matrix Multiply and Accumulate
BMSK Bitfield Mask
BREV Bit Reverse
FLO Find Leading One
IABS Integer Absolute Value
IADD Integer Addition
IADD3 3-input Integer Addition
IADD32I Integer Addition
IDP Integer Dot Product and Accumulate
IDP4A Integer Dot Product and Accumulate
IMAD Integer Multiply And Add
IMMA Integer Matrix Multiply and Accumulate
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
IMUL32I Integer Multiply
ISCADD Scaled Integer Addition
ISCADD32I Scaled Integer Addition
ISETP Integer Compare And Set Predicate
LEA LOAD Effective Address
LOP Logic Operation
LOP3 Logic Operation
LOP32I Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
VABSDIFF Absolute Difference
VABSDIFF4 Absolute Difference

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion
I2IP Integer To Integer Conversion and Packing
FRND Round To Integer

Movement Instructions

Opcode Description
MOV Move
MOV32I Move
MOVM Move Matrix with Transposition or Expansion
PRMT Permute Register Pair
SEL Select Source with Predicate
SGXT Sign Extend
SHFL Warp Wide Register Shuffle

Predicate Instructions

Opcode Description
PLOP3 Predicate Logic Operation
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate Register

Load/Store Instructions

Opcode Description
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDL Load within Local Memory Window
LDS Load within Shared Memory Window
LDSM Load Matrix from Shared Memory with Element Size Expansion
ST Store to Generic Memory
STG Store to Global Memory
STL Store within Local or Shared Window
STS Store within Local or Shared Window
MATCH Match Register Values Across Thread Group
QSPC Query Space
ATOM Atomic Operation on Generic Memory
ATOMS Atomic Operation on Shared Memory
ATOMG Atomic Operation on Global Memory
RED Reduction Operation on Generic Memory
CCTL Cache Control
CCTLL Cache Control
ERRBAR Error Barrier
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Uniform Datapath Instructions

Opcode Description
R2UR Move from Vector Register to a Uniform Register
S2UR Move Special Register to Uniform Register
UBMSK Uniform Bitfield Mask
UBREV Uniform Bit Reverse
UCLEA Load Effective Address for a Constant
UFLO Uniform Find Leading One
UIADD3 Uniform Integer Addition
UIADD3.64 Uniform Integer Addition
UIMAD Uniform Integer Multiplication
UISETP Integer Compare and Set Uniform Predicate
ULDC Load from Constant Memory into a Uniform Register
ULEA Uniform Load Effective Address
ULOP Logic Operation
ULOP3 Logic Operation
ULOP32I Logic Operation
UMOV Uniform Move
UP2UR Uniform Predicate to Uniform Register
UPLOP3 Uniform Predicate Logic Operation
UPOPC Uniform Population Count
UPRMT Uniform Byte Permute
UPSETP Uniform Predicate Logic Operation
UR2UP Uniform Register to Uniform Predicate
USEL Uniform Select
USGXT Uniform Sign Extend
USHF Uniform Funnel Shift
USHL Uniform Left Shift
USHR Uniform Right Shift
VOTEU Voting across SIMD Thread Group with Results in Uniform Destination

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TMML Texture MipMap Level
TXD Texture Fetch With Derivatives
TXQ Texture Query

Surface Instructions

Opcode Description
SUATOM Surface Reduction
SULD Surface Load
SURED Atomic Reduction on surface memory
SUST Surface Store

Control Instructions

Opcode Description
BMOV Move Convergence Barrier State
BPT BreakPoint/Trap
BRA Relative Branch
BREAK Break out of the Specified Convergence Barrier
BRX Relative Branch Indirect
BRXU Relative Branch with Uniform Register Based Offset
BSSY Barrier Set Convergence Synchronization Point
BSYNC Synchronize Threads on a Convergence Barrier
CALL Call Function
EXIT Exit Program
JMP Absolute Jump
JMX Absolute Jump Indirect
JMXU Absolute Jump with Uniform Register Based Offset
KILL Kill Thread
NANOSLEEP Suspend Execution
RET Return From Subroutine
RPCMOV PC Register Move
RTT Return From Trap
WARPSYNC Synchronize Threads in Warp
YIELD Yield Control

Miscellaneous Instructions

Opcode Description
B2R Move Barrier To Register
BAR Barrier Synchronization
CS2R Move Special Register to Register
CSMTEST Clip State Machine Test and Update
DEPBAR Dependency Barrier
GETLMEMBASE Get Local Memory Base Address
LEPC Load Effective PC
NOP No Operation
PMTRIG Performance Monitor Trigger
R2B Move Register to Barrier
S2R Move Special Register to Register
SETCTAID Set CTA ID
SETLMEMBASE Set Local Memory Base Address
VOTE Vote Across SIMD Thread Group
VOTE_VTG Clip State Machine Test and Update

Turing (SM 7.5)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FADD32I FP32 Add
FCHK Floating-point Range Check
FFMA32I FP32 Fused Multiply and Add
FFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMUL FP32 Multiply
FMUL32I FP32 Multiply
FSEL Floating Point Select
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Swizzle Add
MUFU FP32 Multi Function Operation
HADD2 FP16 Add
HADD2_32I FP16 Add
HFMA2 FP16 Fused Mutiply Add
HFMA2_32I FP16 Fused Mutiply Add
HMMA Matrix Multiply and Accumulate
HMUL2 FP16 Multiply
HMUL2_32I FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMUL FP64 Multiply
DSETP FP64 Compare And Set Predicate

Integer Instructions

Opcode Description
BMMA Bit Matrix Multiply and Accumulate
BMSK Bitfield Mask
BREV Bit Reverse
FLO Find Leading One
IABS Integer Absolute Value
IADD Integer Addition
IADD3 3-input Integer Addition
IADD32I Integer Addition
IDP Integer Dot Product and Accumulate
IDP4A Integer Dot Product and Accumulate
IMAD Integer Multiply And Add
IMMA Integer Matrix Multiply and Accumulate
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
IMUL32I Integer Multiply
ISCADD Scaled Integer Addition
ISCADD32I Scaled Integer Addition
ISETP Integer Compare And Set Predicate
LEA LOAD Effective Address
LOP Logic Operation
LOP3 Logic Operation
LOP32I Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
VABSDIFF Absolute Difference
VABSDIFF4 Absolute Difference

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion
I2IP Integer To Integer Conversion and Packing
FRND Round To Integer

Movement Instructions

Opcode Description
MOV Move
MOV32I Move
MOVM Move Matrix with Transposition or Expansion
PRMT Permute Register Pair
SEL Select Source with Predicate
SGXT Sign Extend
SHFL Warp Wide Register Shuffle

Predicate Instructions

Opcode Description
PLOP3 Predicate Logic Operation
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate Register

Load/Store Instructions

Opcode Description
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDL Load within Local Memory Window
LDS Load within Shared Memory Window
LDSM Load Matrix from Shared Memory with Element Size Expansion
ST Store to Generic Memory
STG Store to Global Memory
STL Store to Local Memory
STS Store to Shared Memory
MATCH Match Register Values Across Thread Group
QSPC Query Space
ATOM Atomic Operation on Generic Memory
ATOMS Atomic Operation on Shared Memory
ATOMG Atomic Operation on Global Memory
RED Reduction Operation on Generic Memory
CCTL Cache Control
CCTLL Cache Control
ERRBAR Error Barrier
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Uniform Datapath Instructions

Opcode Description
R2UR Move from Vector Register to a Uniform Register
S2UR Move Special Register to Uniform Register
UBMSK Uniform Bitfield Mask
UBREV Uniform Bit Reverse
UCLEA Load Effective Address for a Constant
UFLO Uniform Find Leading One
UIADD3 Uniform Integer Addition
UIADD3.64 Uniform Integer Addition
UIMAD Uniform Integer Multiplication
UISETP Integer Compare and Set Uniform Predicate
ULDC Load from Constant Memory into a Uniform Register
ULEA Uniform Load Effective Address
ULOP Logic Operation
ULOP3 Logic Operation
ULOP32I Logic Operation
UMOV Uniform Move
UP2UR Uniform Predicate to Uniform Register
UPLOP3 Uniform Predicate Logic Operation
UPOPC Uniform Population Count
UPRMT Uniform Byte Permute
UPSETP Uniform Predicate Logic Operation
UR2UP Uniform Register to Uniform Predicate
USEL Uniform Select
USGXT Uniform Sign Extend
USHF Uniform Funnel Shift
USHL Uniform Left Shift
USHR Uniform Right Shift
VOTEU Voting across SIMD Thread Group with Results in Uniform Destination

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TMML Texture MipMap Level
TXD Texture Fetch With Derivatives
TXQ Texture Query

Surface Instructions

Opcode Description
SUATOM Atomic Op on Surface Memory
SULD Surface Load
SURED Reduction Op on Surface Memory
SUST Surface Store

Control Instructions

Opcode Description
BMOV Move Convergence Barrier State
BPT BreakPoint/Trap
BRA Relative Branch
BREAK Break out of the Specified Convergence Barrier
BRX Relative Branch Indirect
BRXU Relative Branch with Uniform Register Based Offset
BSSY Barrier Set Convergence Synchronization Point
BSYNC Synchronize Threads on a Convergence Barrier
CALL Call Function
EXIT Exit Program
JMP Absolute Jump
JMX Absolute Jump Indirect
JMXU Absolute Jump with Uniform Register Based Offset
KILL Kill Thread
NANOSLEEP Suspend Execution
RET Return From Subroutine
RPCMOV PC Register Move
RTT Return From Trap
WARPSYNC Synchronize Threads in Warp
YIELD Yield Control

Miscellaneous Instructions

Opcode Description
B2R Move Barrier To Register
BAR Barrier Synchronization
CS2R Move Special Register to Register
DEPBAR Dependency Barrier
GETLMEMBASE Get Local Memory Base Address
LEPC Load Effective PC
NOP No Operation
PMTRIG Performance Monitor Trigger
R2B Move Register to Barrier
S2R Move Special Register to Register
SETCTAID Set CTA ID
SETLMEMBASE Set Local Memory Base Address
VOTE Vote Across SIMD Thread Group

Ampere and Ada (SM 8.0 / 8.6 / 8.9)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FADD32I FP32 Add
FCHK Floating-point Range Check
FFMA32I FP32 Fused Multiply and Add
FFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMUL FP32 Multiply
FMUL32I FP32 Multiply
FSEL Floating Point Select
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Swizzle Add
MUFU FP32 Multi Function Operation
HADD2 FP16 Add
HADD2_32I FP16 Add
HFMA2 FP16 Fused Mutiply Add
HFMA2_32I FP16 Fused Mutiply Add
HMMA Matrix Multiply and Accumulate
HMNMX2 FP16 Minimum / Maximum
HMUL2 FP16 Multiply
HMUL2_32I FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMMA Matrix Multiply and Accumulate
DMUL FP64 Multiply
DSETP FP64 Compare And Set Predicate

Integer Instructions

Opcode Description
BMMA Bit Matrix Multiply and Accumulate
BMSK Bitfield Mask
BREV Bit Reverse
FLO Find Leading One
IABS Integer Absolute Value
IADD Integer Addition
IADD3 3-input Integer Addition
IADD32I Integer Addition
IDP Integer Dot Product and Accumulate
IDP4A Integer Dot Product and Accumulate
IMAD Integer Multiply And Add
IMMA Integer Matrix Multiply and Accumulate
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
IMUL32I Integer Multiply
ISCADD Scaled Integer Addition
ISCADD32I Scaled Integer Addition
ISETP Integer Compare And Set Predicate
LEA LOAD Effective Address
LOP Logic Operation
LOP3 Logic Operation
LOP32I Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
VABSDIFF Absolute Difference
VABSDIFF4 Absolute Difference

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion
I2IP Integer To Integer Conversion and Packing
I2FP Integer to FP32 Convert and Pack
F2IP FP32 Down-Convert to Integer and Pack
FRND Round To Integer

Movement Instructions

Opcode Description
MOV Move
MOV32I Move
MOVM Move Matrix with Transposition or Expansion
PRMT Permute Register Pair
SEL Select Source with Predicate
SGXT Sign Extend
SHFL Warp Wide Register Shuffle

Predicate Instructions

Opcode Description
PLOP3 Predicate Logic Operation
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate Register

Load/Store Instructions

Opcode Description
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDGDEPBAR Global Load Dependency Barrier
LDGSTS Asynchronous Global to Shared Memcopy
LDL Load within Local Memory Window
LDS Load within Shared Memory Window
LDSM Load Matrix from Shared Memory with Element Size Expansion
ST Store to Generic Memory
STG Store to Global Memory
STL Store to Local Memory
STS Store to Shared Memory
MATCH Match Register Values Across Thread Group
QSPC Query Space
ATOM Atomic Operation on Generic Memory
ATOMS Atomic Operation on Shared Memory
ATOMG Atomic Operation on Global Memory
RED Reduction Operation on Generic Memory
CCTL Cache Control
CCTLL Cache Control
ERRBAR Error Barrier
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Uniform Datapath Instructions

Opcode Description
R2UR Move from Vector Register to a Uniform Register
REDUX Reduction of a Vector Register into a Uniform Register
S2UR Move Special Register to Uniform Register
UBMSK Uniform Bitfield Mask
UBREV Uniform Bit Reverse
UCLEA Load Effective Address for a Constant
UF2FP Uniform FP32 Down-convert and Pack
UFLO Uniform Find Leading One
UIADD3 Uniform Integer Addition
UIADD3.64 Uniform Integer Addition
UIMAD Uniform Integer Multiplication
UISETP Integer Compare and Set Uniform Predicate
ULDC Load from Constant Memory into a Uniform Register
ULEA Uniform Load Effective Address
ULOP Logic Operation
ULOP3 Logic Operation
ULOP32I Logic Operation
UMOV Uniform Move
UP2UR Uniform Predicate to Uniform Register
UPLOP3 Uniform Predicate Logic Operation
UPOPC Uniform Population Count
UPRMT Uniform Byte Permute
UPSETP Uniform Predicate Logic Operation
UR2UP Uniform Register to Uniform Predicate
USEL Uniform Select
USGXT Uniform Sign Extend
USHF Uniform Funnel Shift
USHL Uniform Left Shift
USHR Uniform Right Shift
VOTEU Voting across SIMD Thread Group with Results in Uniform Destination

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TMML Texture MipMap Level
TXD Texture Fetch With Derivatives
TXQ Texture Query

Surface Instructions

Opcode Description
SUATOM Atomic Op on Surface Memory
SULD Surface Load
SURED Reduction Op on Surface Memory
SUST Surface Store

Control Instructions

Opcode Description
BMOV Move Convergence Barrier State
BPT BreakPoint/Trap
BRA Relative Branch
BREAK Break out of the Specified Convergence Barrier
BRX Relative Branch Indirect
BRXU Relative Branch with Uniform Register Based Offset
BSSY Barrier Set Convergence Synchronization Point
BSYNC Synchronize Threads on a Convergence Barrier
CALL Call Function
EXIT Exit Program
JMP Absolute Jump
JMX Absolute Jump Indirect
JMXU Absolute Jump with Uniform Register Based Offset
KILL Kill Thread
NANOSLEEP Suspend Execution
RET Return From Subroutine
RPCMOV PC Register Move
WARPSYNC Synchronize Threads in Warp
YIELD Yield Control

Miscellaneous Instructions

Opcode Description
B2R Move Barrier To Register
BAR Barrier Synchronization
CS2R Move Special Register to Register
DEPBAR Dependency Barrier
GETLMEMBASE Get Local Memory Base Address
LEPC Load Effective PC
NOP No Operation
PMTRIG Performance Monitor Trigger
S2R Move Special Register to Register
SETCTAID Set CTA ID
SETLMEMBASE Set Local Memory Base Address
VOTE Vote Across SIMD Thread Group

Hopper (SM 9.0)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FADD32I FP32 Add
FCHK Floating-point Range Check
FFMA32I FP32 Fused Multiply and Add
FFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMUL FP32 Multiply
FMUL32I FP32 Multiply
FSEL Floating Point Select
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Swizzle Add
MUFU FP32 Multi Function Operation
HADD2 FP16 Add
HADD2_32I FP16 Add
HFMA2 FP16 Fused Mutiply Add
HFMA2_32I FP16 Fused Mutiply Add
HMMA Matrix Multiply and Accumulate
HMNMX2 FP16 Minimum / Maximum
HMUL2 FP16 Multiply
HMUL2_32I FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMMA Matrix Multiply and Accumulate
DMUL FP64 Multiply
DSETP FP64 Compare And Set Predicate

Integer Instructions

Opcode Description
BMMA Bit Matrix Multiply and Accumulate
BMSK Bitfield Mask
BREV Bit Reverse
FLO Find Leading One
IABS Integer Absolute Value
IADD Integer Addition
IADD3 3-input Integer Addition
IADD32I Integer Addition
IDP Integer Dot Product and Accumulate
IDP4A Integer Dot Product and Accumulate
IMAD Integer Multiply And Add
IMMA Integer Matrix Multiply and Accumulate
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
IMUL32I Integer Multiply
ISCADD Scaled Integer Addition
ISCADD32I Scaled Integer Addition
ISETP Integer Compare And Set Predicate
LEA LOAD Effective Address
LOP Logic Operation
LOP3 Logic Operation
LOP32I Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
VABSDIFF Absolute Difference
VABSDIFF4 Absolute Difference
VHMNMX SIMD FP16 3-Input Minimum / Maximum
VIADD SIMD Integer Addition
VIADDMNMX SIMD Integer Addition and Fused Min/Max Comparison
VIMNMX SIMD Integer Minimum / Maximum
VIMNMX3 SIMD Integer 3-Input Minimum / Maximum

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion
I2IP Integer To Integer Conversion and Packing
I2FP Integer to FP32 Convert and Pack
F2IP FP32 Down-Convert to Integer and Pack
FRND Round To Integer

Movement Instructions

Opcode Description
MOV Move
MOV32I Move
MOVM Move Matrix with Transposition or Expansion
PRMT Permute Register Pair
SEL Select Source with Predicate
SGXT Sign Extend
SHFL Warp Wide Register Shuffle

Predicate Instructions

Opcode Description
PLOP3 Predicate Logic Operation
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate Register

Load/Store Instructions

Opcode Description
FENCE Memory Visibility Guarantee for Shared or Global Memory
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDGDEPBAR Global Load Dependency Barrier
LDGMC Reducing Load
LDGSTS Asynchronous Global to Shared Memcopy
LDL Load within Local Memory Window
LDS Load within Shared Memory Window
LDSM Load Matrix from Shared Memory with Element Size Expansion
STSM Store Matrix to Shared Memory
ST Store to Generic Memory
STG Store to Global Memory
STL Store to Local Memory
STS Store to Shared Memory
STAS Asynchronous Store to Distributed Shared Memory With Explicit Synchronization
SYNCS Sync Unit
MATCH Match Register Values Across Thread Group
QSPC Query Space
ATOM Atomic Operation on Generic Memory
ATOMS Atomic Operation on Shared Memory
ATOMG Atomic Operation on Global Memory
REDAS Asynchronous Reduction on Distributed Shared Memory With Explicit Synchronization
REDG Reduction Operation on Generic Memory
CCTL Cache Control
CCTLL Cache Control
ERRBAR Error Barrier
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Uniform Datapath Instructions

Opcode Description
R2UR Move from Vector Register to a Uniform Register
REDUX Reduction of a Vector Register into a Uniform Register
S2UR Move Special Register to Uniform Register
UBMSK Uniform Bitfield Mask
UBREV Uniform Bit Reverse
UCGABAR_ARV CGA Barrier Synchronization
UCGABAR_WAIT CGA Barrier Synchronization
UCLEA Load Effective Address for a Constant
UF2FP Uniform FP32 Down-convert and Pack
UFLO Uniform Find Leading One
UIADD3 Uniform Integer Addition
UIADD3.64 Uniform Integer Addition
UIMAD Uniform Integer Multiplication
UISETP Integer Compare and Set Uniform Predicate
ULDC Load from Constant Memory into a Uniform Register
ULEA Uniform Load Effective Address
ULEPC Uniform Load Effective PC
ULOP Logic Operation
ULOP3 Logic Operation
ULOP32I Logic Operation
UMOV Uniform Move
UP2UR Uniform Predicate to Uniform Register
UPLOP3 Uniform Predicate Logic Operation
UPOPC Uniform Population Count
UPRMT Uniform Byte Permute
UPSETP Uniform Predicate Logic Operation
UR2UP Uniform Register to Uniform Predicate
USEL Uniform Select
USETMAXREG Release, Deallocate and Allocate Registers
USGXT Uniform Sign Extend
USHF Uniform Funnel Shift
USHL Uniform Left Shift
USHR Uniform Right Shift
VOTEU Voting across SIMD Thread Group with Results in Uniform Destination

Warpgroup Instructions

Opcode Description
BGMMA Bit Matrix Multiply and Accumulate Across Warps
HGMMA Matrix Multiply and Accumulate Across a Warpgroup
IGMMA Integer Matrix Multiply and Accumulate Across a Warpgroup
QGMMA FP8 Matrix Multiply and Accumulate Across a Warpgroup
WARPGROUP Warpgroup Synchronization
WARPGROUPSET Set Warpgroup Counters

Tensor Memory Access Instructions

Opcode Description
UBLKCP Bulk Data Copy
UBLKPF Bulk Data Prefetch
UBLKRED Bulk Data Copy from Shared Memory with Reduction
UTMACCTL TMA Cache Control
UTMACMDFLUSH TMA Command Flush
UTMALDG Tensor Load from Global to Shared Memory
UTMAPF Tensor Prefetch
UTMAREDG Tensor Store from Shared to Global Memory with Reduction
UTMASTG Tensor Store from Shared to Global Memory

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TMML Texture MipMap Level
TXD Texture Fetch With Derivatives
TXQ Texture Query

Surface Instructions

Opcode Description
SUATOM Atomic Op on Surface Memory
SULD Surface Load
SURED Reduction Op on Surface Memory
SUST Surface Store

Control Instructions

Opcode Description
ACQBULK Wait for Bulk Release Status Warp State
BMOV Move Convergence Barrier State
BPT BreakPoint/Trap
BRA Relative Branch
BREAK Break out of the Specified Convergence Barrier
BRX Relative Branch Indirect
BRXU Relative Branch with Uniform Register Based Offset
BSSY Barrier Set Convergence Synchronization Point
BSYNC Synchronize Threads on a Convergence Barrier
CALL Call Function
CGAERRBAR CGA Error Barrier
ELECT Elect a Leader Thread
ENDCOLLECTIVE Reset the MCOLLECTIVE mask
EXIT Exit Program
JMP Absolute Jump
JMX Absolute Jump Indirect
JMXU Absolute Jump with Uniform Register Based Offset
KILL Kill Thread
NANOSLEEP Suspend Execution
PREEXIT Dependent Task Launch Hint
RET Return From Subroutine
RPCMOV PC Register Move
WARPSYNC Synchronize Threads in Warp
YIELD Yield Control

Miscellaneous Instructions

Opcode Description
B2R Move Barrier To Register
BAR Barrier Synchronization
CS2R Move Special Register to Register
DEPBAR Dependency Barrier
GETLMEMBASE Get Local Memory Base Address
LEPC Load Effective PC
NOP No Operation
PMTRIG Performance Monitor Trigger
S2R Move Special Register to Register
SETCTAID Set CTA ID
SETLMEMBASE Set Local Memory Base Address
VOTE Vote Across SIMT Thread Group

Blackwell (SM 10.x / 12.x)

Floating Point Instructions

Opcode Description
FADD FP32 Add
FADD2 FP32 Add
FADD32I FP32 Add
FCHK Floating-point Range Check
FFMA32I FP32 Fused Multiply and Add
FFMA FP32 Fused Multiply and Add
FFMA2 FP32 Fused Multiply and Add
FHADD FP32 Addition
FHFMA FP32 Fused Multiply and Add
FMNMX FP32 Minimum/Maximum
FMNMX3 3-Input Floating-point Minimum / Maximum
FMUL FP32 Multiply
FMUL2 FP32 Multiply
FMUL32I FP32 Multiply
FSEL Floating Point Select
FSET FP32 Compare And Set
FSETP FP32 Compare And Set Predicate
FSWZADD FP32 Swizzle Add
MUFU FP32 Multi Function Operation
HADD2 FP16 Add
HADD2_32I FP16 Add
HFMA2 FP16 Fused Mutiply Add
HFMA2_32I FP16 Fused Mutiply Add
HMMA Matrix Multiply and Accumulate
HMNMX2 FP16 Minimum / Maximum
HMUL2 FP16 Multiply
HMUL2_32I FP16 Multiply
HSET2 FP16 Compare And Set
HSETP2 FP16 Compare And Set Predicate
DADD FP64 Add
DFMA FP64 Fused Mutiply Add
DMMA Matrix Multiply and Accumulate
DMUL FP64 Multiply
DSETP FP64 Compare And Set Predicate
OMMA FP4 Matrix Multiply and Accumulate Across a Warp
QMMA FP8 Matrix Multiply and Accumulate Across a Warp

Integer Instructions

Opcode Description
BMSK Bitfield Mask
BREV Bit Reverse
FLO Find Leading One
IABS Integer Absolute Value
IADD Integer Addition
IADD3 3-input Integer Addition
IADD32I Integer Addition
IDP Integer Dot Product and Accumulate
IDP4A Integer Dot Product and Accumulate
IMAD Integer Multiply And Add
IMMA Integer Matrix Multiply and Accumulate
IMNMX Integer Minimum/Maximum
IMUL Integer Multiply
IMUL32I Integer Multiply
ISCADD Scaled Integer Addition
ISCADD32I Scaled Integer Addition
ISETP Integer Compare And Set Predicate
LEA LOAD Effective Address
LOP Logic Operation
LOP3 Logic Operation
LOP32I Logic Operation
POPC Population count
SHF Funnel Shift
SHL Shift Left
SHR Shift Right
VABSDIFF Absolute Difference
VABSDIFF4 Absolute Difference
VHMNMX SIMD FP16 3-Input Minimum / Maximum
VIADD SIMD Integer Addition
VIADDMNMX SIMD Integer Addition and Fused Min/Max Comparison
VIMNMX SIMD Integer Minimum / Maximum
VIMNMX3 SIMD Integer 3-Input Minimum / Maximum

Conversion Instructions

Opcode Description
F2F Floating Point To Floating Point Conversion
F2I Floating Point To Integer Conversion
I2F Integer To Floating Point Conversion
I2I Integer To Integer Conversion
I2IP Integer To Integer Conversion and Packing
I2FP Integer to FP32 Convert and Pack
F2IP FP32 Down-Convert to Integer and Pack
FRND Round To Integer

Movement Instructions

Opcode Description
MOV Move
MOV32I Move
MOVM Move Matrix with Transposition or Expansion
PRMT Permute Register Pair
SEL Select Source with Predicate
SGXT Sign Extend
SHFL Warp Wide Register Shuffle

Predicate Instructions

Opcode Description
PLOP3 Predicate Logic Operation
PSETP Combine Predicates and Set Predicate
P2R Move Predicate Register To Register
R2P Move Register To Predicate Register

Load/Store Instructions

Opcode Description
FENCE Memory Visibility Guarantee for Shared or Global Memory
LD Load from generic Memory
LDC Load Constant
LDG Load from Global Memory
LDGDEPBAR Global Load Dependency Barrier
LDGMC Reducing Load
LDGSTS Asynchronous Global to Shared Memcopy
LDL Load within Local Memory Window
LDS Load within Shared Memory Window
LDSM Load Matrix from Shared Memory with Element Size Expansion
STSM Store Matrix to Shared Memory
ST Store to Generic Memory
STG Store to Global Memory
STL Store to Local Memory
STS Store to Shared Memory
STAS Asynchronous Store to Distributed Shared Memory With Explicit Synchronization
SYNCS Sync Unit
MATCH Match Register Values Across Thread Group
QSPC Query Space
ATOM Atomic Operation on Generic Memory
ATOMS Atomic Operation on Shared Memory
ATOMG Atomic Operation on Global Memory
REDAS Asynchronous Reduction on Distributed Shared Memory With Explicit Synchronization
REDG Reduction Operation on Generic Memory
CCTL Cache Control
CCTLL Cache Control
ERRBAR Error Barrier
MEMBAR Memory Barrier
CCTLT Texture Cache Control

Uniform Datapath Instructions

Opcode Description
CREDUX Coupled Reduction of a Vector Register into a Uniform Register
CS2UR Load a Value from Constant Memory into a Uniform Register
LDCU Load a Value from Constant Memory into a Uniform Register
R2UR Move from Vector Register to a Uniform Register
REDUX Reduction of a Vector Register into a Uniform Register
S2UR Move Special Register to Uniform Register
UBMSK Uniform Bitfield Mask
UBREV Uniform Bit Reverse
UCGABAR_ARV CGA Barrier Synchronization
UCGABAR_WAIT CGA Barrier Synchronization
UCLEA Load Effective Address for a Constant
UFADD Uniform Uniform FP32 Addition
UF2F Uniform Float-to-Float Conversion
UF2FP Uniform FP32 Down-convert and Pack
UF2I Uniform Float-to-Integer Conversion
UF2IP Uniform FP32 Down-Convert to Integer and Pack
UFFMA Uniform FP32 Fused Multiply-Add
UFLO Uniform Find Leading One
UFMNMX Uniform Floating-point Minimum / Maximum
UFMUL Uniform FP32 Multiply
UFRND Uniform Round to Integer
UFSEL Uniform Floating-Point Select
UFSET Uniform Floating-Point Compare and Set
UFSETP Uniform Floating-Point Compare and Set Predicate
UI2F Uniform Integer to Float conversion
UI2FP Uniform Integer to FP32 Convert and Pack
UI2I Uniform Saturating Integer-to-Integer Conversion
UI2IP Uniform Dual Saturating Integer-to-Integer Conversion and Packing
UIABS Uniform Integer Absolute Value
UIMNMX Uniform Integer Minimum / Maximum
UIADD3 Uniform Integer Addition
UIADD3.64 Uniform Integer Addition
UIMAD Uniform Integer Multiplication
UISETP Uniform Integer Compare and Set Uniform Predicate
ULEA Uniform Load Effective Address
ULEPC Uniform Load Effective PC
ULOP Uniform Logic Operation
ULOP3 Uniform Logic Operation
ULOP32I Uniform Logic Operation
UMOV Uniform Move
UP2UR Uniform Predicate to Uniform Register
UPLOP3 Uniform Predicate Logic Operation
UPOPC Uniform Population Count
UPRMT Uniform Byte Permute
UPSETP Uniform Predicate Logic Operation
UR2UP Uniform Register to Uniform Predicate
USEL Uniform Select
USETMAXREG Release, Deallocate and Allocate Registers
USGXT Uniform Sign Extend
USHF Uniform Funnel Shift
USHL Uniform Left Shift
USHR Uniform Right Shift
UGETNEXTWORKID Uniform Get Next Work ID
UMEMSETS Initialize Shared Memory
UREDGR Uniform Reduction on Global Memory with Release
USTGR Uniform Store to Global Memory with Release
UVIADD Uniform SIMD Integer Addition
UVIMNMX Uniform SIMD Integer Minimum / Maximum
UVIRTCOUNT Virtual Resource Management
VOTEU Voting across SIMD Thread Group with Results in Uniform Destination

Tensor Memory Access Instructions

Opcode Description
UBLKCP Bulk Data Copy
UBLKPF Bulk Data Prefetch
UBLKRED Bulk Data Copy from Shared Memory with Reduction
UTMACCTL TMA Cache Control
UTMACMDFLUSH TMA Command Flush
UTMALDG Tensor Load from Global to Shared Memory
UTMAPF Tensor Prefetch
UTMAREDG Tensor Store from Shared to Global Memory with Reduction
UTMASTG Tensor Store from Shared to Global Memory

Tensor Core Memory Instructions

Opcode Description
LDT Load Matrix from Tensor Memory to Register File
LDTM Load Matrix from Tensor Memory to Register File
STT Store Matrix to Tensor Memory from Register File
STTM Store Matrix to Tensor Memory from Register File
UTCATOMSWS Perform Atomic operation on SW State Register
UTCBAR Tensor Core Barrier
UTCCP Asynchonous data copy from Shared Memory to Tensor Memory
UTCHMMA Uniform Matrix Multiply and Accumulate
UTCIMMA Uniform Matrix Multiply and Accumulate
UTCOMMA Uniform Matrix Multiply and Accumulate
UTCQMMA Uniform Matrix Multiply and Accumulate
UTCSHIFT Shift elements in Tensor Memory

Texture Instructions

Opcode Description
TEX Texture Fetch
TLD Texture Load
TLD4 Texture Load 4
TMML Texture MipMap Level
TXD Texture Fetch With Derivatives
TXQ Texture Query

Surface Instructions

Opcode Description
SUATOM Atomic Op on Surface Memory
SULD Surface Load
SURED Reduction Op on Surface Memory
SUST Surface Store

Control Instructions

Opcode Description
ACQBULK Wait for Bulk Release Status Warp State
ACQSHMINIT Wait for Shared Memory Initialization Release Status Warp State
BMOV Move Convergence Barrier State
BPT BreakPoint/Trap
BRA Relative Branch
BREAK Break out of the Specified Convergence Barrier
BRX Relative Branch Indirect
BRXU Relative Branch with Uniform Register Based Offset
BSSY Barrier Set Convergence Synchronization Point
BSYNC Synchronize Threads on a Convergence Barrier
CALL Call Function
CGAERRBAR CGA Error Barrier
ELECT Elect a Leader Thread
ENDCOLLECTIVE Reset the MCOLLECTIVE mask
EXIT Exit Program
JMP Absolute Jump
JMX Absolute Jump Indirect
JMXU Absolute Jump with Uniform Register Based Offset
KILL Kill Thread
NANOSLEEP Suspend Execution
PREEXIT Dependent Task Launch Hint
RET Return From Subroutine
RPCMOV PC Register Move
WARPSYNC Synchronize Threads in Warp
YIELD Yield Control

Miscellaneous Instructions

Opcode Description
B2R Move Barrier To Register
BAR Barrier Synchronization
CS2R Move Special Register to Register
DEPBAR Dependency Barrier
GETLMEMBASE Get Local Memory Base Address
LEPC Load Effective PC
NOP No Operation
PMTRIG Performance Monitor Trigger
S2R Move Special Register to Register
SETCTAID Set CTA ID
SETLMEMBASE Set Local Memory Base Address
VOTE Vote Across SIMT Thread Group