Since the first CUDA-capable hardware shipped in 2006, NVIDIA has developed eleven major architectures – Tesla, Fermi, Kepler, Maxwell, Pascal, Volta, Turing, Ampere, Ada Lovelace, Hopper, and Blackwell (Table 8-1) – and the native instruction set has grown with each of them, from a few dozen operations to well over a hundred per generation. New instructions appear both between architectures and, occasionally, within a family as NVIDIA refreshes its products. Global atomic operations, for example, were absent from the very first Tesla-class processor (the G80, which shipped in 2006 as the GeForce 8800 GTX) but present in every Tesla GPU thereafter: querying the compute capability returns 1.0 for the G80 and 1.1 or greater for the rest, so an application can use global atomics whenever the reported version is at least 1.1.
The native instructions are called SASS. CUDA does not
document SASS in the way it documents PTX: NVIDIA publishes no encodings
or semantics for it, reserves the right to change it from one
architecture to the next, and offers no guarantee of binary
compatibility across generations – which is precisely what lets the
instruction set evolve as freely as it has. What NVIDIA does publish, in
the CUDA
Binary Utilities documentation, is a per-architecture list of the
instruction mnemonics that cuobjdump and nvdisasm emit when
disassembling a binary, each with a one-line description. That reference
is the authoritative, living catalog of the SASS instruction set for
every architecture from Turing (SM 7.5) forward. The tables that follow
give the SASS instruction set of every architecture from Tesla through
Blackwell – from that reference for Turing onward, from the edition
archived with CUDA 10.0 for Fermi through Volta, and from the first
edition of this book for the Tesla generation, which predates it.
The instruction sets below give the SASS mnemonics for every CUDA architecture, grouped by instruction category, with the opcode as it appears in disassembly. Fermi through Volta are drawn from the CUDA Binary Utilities reference archived with CUDA 10.0; Turing through Blackwell from the current edition; and the Tesla set, which predates both, from the first edition of this book, whose middle column gives the first SM version to support each instruction.
Floating Point
| Opcode | SM | Description |
|---|---|---|
COS |
1.0 | Cosine |
DADD |
1.3 | Double-precision floating point add |
DFMA |
1.3 | Double-precision floating point fused multiply-add |
DMAX |
1.3 | Double-precision floating point maximum |
DMIN |
1.3 | Double-precision floating point minimum |
DMUL |
1.3 | Double-precision floating point multiply |
DSET |
1.3 | Double-precision floating point condition set |
EX2 |
1.0 | Exponential (base 2) |
FADD/ |
1.0 | Single-precision floating point add |
FCMP |
1.0 | Single-precision floating point compare |
FMAD/ |
1.0 | Single-precision floating point multiply-add* |
FMAX |
1.0 | Single-precision floating point maximum |
FMIN |
1.0 | Single-precision floating point minimum |
FMUL/ |
1.0 | Single-precision floating point multiply |
FSET |
1.0 | Single-precision floating point conditional set |
LG2 |
1.0 | Single-precision floating point logarithm (base 2) |
RCP |
1.0 | Single-precision floating point reciprocal |
RRO |
1.0 | Range reduction operator (used before SIN/COS) |
RSQ |
1.0 | Reciprocal square root |
SIN |
1.0 | Sine |
Flow Control
| Opcode | SM | Description |
|---|---|---|
BAR |
1.0 | Barrier synchronization/ __syncthreads() |
BRA |
1.0 | Conditional branch |
BRK |
1.0 | Conditional break from loop |
BRX |
1.0 | Fetch an address from constant memory and branch to it |
C2R |
1.0 | Condition code to data register |
CAL |
1.0 | Unconditional subroutine call |
RET |
1.0 | Conditional return from subroutine |
SSY |
1.0 | Set synchronization point; used before potentially divergent instructions |
Data Conversion
| Opcode | SM | Description |
|---|---|---|
F2F |
1.0 | Copy floating point value with conversion to floating point |
F2I |
1.0 | Copy floating point value with conversion to integer |
I2F |
1.0 | Copy integer value to floating-point with conversion |
I2I |
1.0 | Copy integer value to integer with conversion |
Integer
| Opcode | SM | Description |
|---|---|---|
IADD/ |
1.0 | Integer addition |
IMAD/ |
1.0 | Integer multiply-add |
IMAX |
1.0 | Integer maximum |
IMIN |
1.0 | Integer minimum |
IMUL/ |
1.0 | Integer multiply |
ISAD/ |
1.0 | Integer sum of absolute difference |
ISET |
1.0 | Integer conditional set |
SHL |
1.0 | Shift left |
SHR |
1.0 | Shift right |
Memory Operations
| Opcode | SM | Description |
|---|---|---|
A2R |
1.0 | Move address register to data register |
ADA |
1.0 | Add immediate to address register |
G2R |
1.0 | Move from shared memory to register. A .LCK suffix indicates that the bank is locked until an R2G.UNL has been performed; this is used to implement shared memory atomics. |
GATOM.IADD/ |
1.2 | Global memory atomic operations; performs an atomic operation and returns the original value. |
GLD |
1.0 | Load from global memory |
GRED.IADD/ |
1.2 | Global memory reduction operations; performs an atomic operation with no return value. |
GST |
1.0 | Store to global memory |
LLD |
1.0 | Load from local memory |
LST |
1.0 | Store to local memory |
LOP |
1.0 | Logical operation (AND/OR/XOR) |
MOV/ |
1.0 | Move source to destination |
MVC |
1.0 | Move from constant memory |
MVI |
1.0 | Move immediate |
R2A |
1.0 | Move register to address register |
R2C |
1.0 | Move data register to condition code |
R2G |
1.0 | Store to shared memory. When used with the .UNL suffix, releases a previously held lock on that shared memory bank. |
Miscellaneous
| Opcode | SM | Description |
|---|---|---|
NOP |
1.0 | No operation |
TEX/ |
1.0 | Texture fetch |
VOTE |
1.2 | Warp-vote primitive. |
S2R |
1.0 | Move special register (e.g., thread ID) to register |
S2R |
1.0 | Move special register (e.g., thread ID) to register |
Floating Point Instructions
| Opcode | Description |
|---|---|
FFMA |
FP32 Fused Multiply Add |
FADD |
FP32 Add |
FCMP |
FP32 Compare |
FMUL |
FP32 Multiply |
FMNMX |
FP32 Minimum/Maximum |
FSWZ |
FP32 Swizzle |
FSET |
FP32 Set |
FSETP |
FP32 Set Predicate |
FCHK |
FP32 Division Test |
RRO |
FP Range Reduction Operator |
MUFU |
FP Multi-Function Operator |
DFMA |
FP64 Fused Multiply Add |
DADD |
FP64 Add |
DMUL |
FP64 Multiply |
DMNMX |
FP64 Minimum/Maximum |
DSET |
FP64 Set |
DSETP |
FP64 Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
IMAD |
Integer Multiply Add |
IMADSP |
Integer Extract Multiply Add |
IMUL |
Integer Multiply |
IADD |
Integer Add |
ISCADD |
Integer Scaled Add |
ISAD |
Integer Sum Of Abs Diff |
IMNMX |
Integer Minimum/Maximum |
BFE |
Integer Bit Field Extract |
BFI |
Integer Bit Field Insert |
SHR |
Integer Shift Right |
SHL |
Integer Shift Left |
SHF |
Integer Funnel Shift |
LOP |
Integer Logic Op |
FLO |
Integer Find Leading One |
ISET |
Integer Set |
ISETP |
Integer Set Predicate |
ICMP |
Integer Compare and Select |
POPC |
Population Count |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Float to Float |
F2I |
Float to Integer |
I2F |
Integer to Float |
I2I |
Integer to Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
SEL |
Conditional Select/Move |
PRMT |
Permute |
SHFL |
Warp Shuffle |
Predicate/CC Instructions
| Opcode | Description |
|---|---|
P2R |
Predicate to Register |
R2P |
Register to Predicate |
CSET |
CC Set |
CSETP |
CC Set Predicate |
PSET |
Predicate Set |
PSETP |
Predicate Set Predicate |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 Texels |
TXQ |
Texture Query |
Compute Load/Store Instructions
| Opcode | Description |
|---|---|
LDC |
Load from Constant |
LD |
Load from Memory |
LDG |
Non-coherent Global Memory Load |
LDL |
Load from Local Memory |
LDS |
Load from Shared Memory |
LDSLK |
Load from Shared Memory and Lock |
ST |
Store to Memory |
STL |
Store to Local Memory |
STS |
Store to Shared Memory |
STSCUL |
Store to Shared Memory Conditionally and Unlock |
ATOM |
Atomic Memory Operation |
RED |
Atomic Memory Reduction Operation |
CCTL |
Cache Control |
CCTLL |
Cache Control (Local) |
MEMBAR |
Memory Barrier |
Surface Memory Instructions
| Opcode | Description |
|---|---|
SUCLAMP |
Surface Clamp |
SUBFM |
Surface Bit Field Merge |
SUEAU |
Surface Effective Address |
SULDGA |
Surface Load Generic Address |
SUSTGA |
Surface Store Generic Address |
Control Instructions
| Opcode | Description |
|---|---|
BRA |
Branch to Relative Address |
BRX |
Branch to Relative Indexed Address |
JMP |
Jump to Absolute Address |
JMX |
Jump to Absolute Indexed Address |
CAL |
Call to Relative Address |
JCAL |
Call to Absolute Address |
RET |
Return from Call |
BRK |
Break from Loop |
CONT |
Continue in Loop |
SSY |
Set Sync Relative Address |
PBK |
Pre-Break Relative Address |
PCNT |
Pre-Continue Relative Address |
PRET |
Pre-Return Relative Address |
BPT |
Breakpoint/Trap |
EXIT |
Exit Program |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
NOP |
No Operation |
S2R |
Special Register to Register |
B2R |
Barrier to Register |
BAR |
Barrier Synchronization |
VOTE |
Query condition across threads |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FCHK |
Single Precision FP Divide Range Check |
FCMP |
FP32 Compare to Zero and Select Source |
FFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMUL |
FP32 Multiply |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Add used for FSWZ emulation |
MUFU |
Multi Function Operation |
RRO |
Range Reduction Operator FP |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMNMX |
FP64 Minimum/Maximum |
DMUL |
FP64 Multiply |
DSET |
FP64 Compare And Set |
DSETP |
FP64 Compare And Set Predicate |
HADD2 |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HMUL2 |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
BFE |
Bit Field Extract |
BFI |
Bit Field Insert |
FLO |
Find Leading One |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
ICMP |
Integer Compare to Zero and Select Source |
IMAD |
Integer Multiply And Add |
IMADSP |
Extracted Integer Multiply And Add. |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISET |
Integer Compare And Set |
ISETP |
Integer Compare And Set Predicate |
LEA |
Compute Effective Address |
LOP |
Logic Operation |
LOP3 |
3-input Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
XMAD |
Integer Short Multiply Add |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SHFL |
Warp Wide Register Shuffle |
Predicate/CC Instructions
| Opcode | Description |
|---|---|
CSET |
Test Condition Code And Set |
CSETP |
Test Condition Code and Set Predicate |
PSET |
Combine Predicates and Set |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate/CC Register |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TXQ |
Texture Query |
TEXS |
Texture Fetch with scalar/non-vec4 source/destinations |
TLD4S |
Texture Load 4 with scalar/non-vec4 source/destinations |
TLDS |
Texture Load with scalar/non-vec4 source/destinations |
Compute Load/Store Instructions
| Opcode | Description |
|---|---|
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDL |
Load within Local Memory Window |
LDS |
Local within Shared Memory Window |
ST |
Store to generic Memory |
STG |
Store to global Memory |
STL |
Store within Local or Shared Window |
STS |
Store within Local or Shared Window |
ATOM |
Atomic Operation on generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
RED |
Reduction Operation on generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Surface Memory Instructions
| Opcode | Description |
|---|---|
SUATOM |
Surface Reduction |
SULD |
Surface Load |
SURED |
Atomic Reduction on surface memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
BRA |
Relative Branch |
BRX |
Relative Branch Indirect |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
SSY |
Set Synchronization Point |
SYNC |
Converge threads after conditional branch |
CAL |
Relative Call |
JCAL |
Absolute Call |
PRET |
Pre-Return From Subroutine |
RET |
Return From Subroutine |
BRK |
Break |
PBK |
Pre-Break |
CONT |
Continue |
PCNT |
Pre-continue |
EXIT |
Exit Program |
PEXIT |
Pre-Exit |
BPT |
BreakPoint/Trap |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
NOP |
No Operation |
CS2R |
Move Special Register to Register |
S2R |
Move Special Register to Register |
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
R2B |
Move Register to Barrier |
VOTE |
Vote Across SIMD Thread Group |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FADD32I |
FP32 Add |
FCHK |
Floating-point Range Check |
FFMA32I |
FP32 Fused Multiply and Add |
FFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMUL |
FP32 Multiply |
FMUL32I |
FP32 Multiply |
FSEL |
Floating Point Select |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Swizzle Add |
MUFU |
FP32 Multi Function Operation |
HADD2 |
FP16 Add |
HADD2_32I |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HFMA2_32I |
FP16 Fused Mutiply Add |
HMMA |
Matrix Multiply and Accumulate |
HMUL2 |
FP16 Multiply |
HMUL2_32I |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMUL |
FP64 Multiply |
DSETP |
FP64 Compare And Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
BMSK |
Bitfield Mask |
BREV |
Bit Reverse |
FLO |
Find Leading One |
IABS |
Integer Absolute Value |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
IADD32I |
Integer Addition |
IDP |
Integer Dot Product and Accumulate |
IDP4A |
Integer Dot Product and Accumulate |
IMAD |
Integer Multiply And Add |
IMMA |
Integer Matrix Multiply and Accumulate |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
IMUL32I |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISCADD32I |
Scaled Integer Addition |
ISETP |
Integer Compare And Set Predicate |
LEA |
LOAD Effective Address |
LOP |
Logic Operation |
LOP3 |
Logic Operation |
LOP32I |
Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
VABSDIFF |
Absolute Difference |
VABSDIFF4 |
Absolute Difference |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
I2IP |
Integer To Integer Conversion and Packing |
FRND |
Round To Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
MOV32I |
Move |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SGXT |
Sign Extend |
SHFL |
Warp Wide Register Shuffle |
Predicate Instructions
| Opcode | Description |
|---|---|
PLOP3 |
Predicate Logic Operation |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate Register |
Load/Store Instructions
| Opcode | Description |
|---|---|
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDL |
Load within Local Memory Window |
LDS |
Load within Shared Memory Window |
ST |
Store to Generic Memory |
STG |
Store to Global Memory |
STL |
Store within Local or Shared Window |
STS |
Store within Local or Shared Window |
MATCH |
Match Register Values Across Thread Group |
QSPC |
Query Space |
ATOM |
Atomic Operation on Generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
ATOMG |
Atomic Operation on Global Memory |
RED |
Reduction Operation on Generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
ERRBAR |
Error Barrier |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TMML |
Texture MipMap Level |
TXD |
Texture Fetch With Derivatives |
TXQ |
Texture Query |
Surface Instructions
| Opcode | Description |
|---|---|
SUATOM |
Surface Reduction |
SULD |
Surface Load |
SURED |
Atomic Reduction on surface memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
BMOV |
Move Convergence Barrier State |
BPT |
BreakPoint/Trap |
BRA |
Relative Branch |
BREAK |
Break out of the Specified Convergence Barrier |
BRX |
Relative Branch Indirect |
BSSY |
Barrier Set Convergence Synchronization Point |
BSYNC |
Synchronize Threads on a Convergence Barrier |
CALL |
Call Function |
EXIT |
Exit Program |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
KILL |
Kill Thread |
NANOSLEEP |
Suspend Execution |
RET |
Return From Subroutine |
RPCMOV |
PC Register Move |
RTT |
Return From Trap |
WARPSYNC |
Synchronize Threads in Warp |
YIELD |
Yield Control |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
CS2R |
Move Special Register to Register |
CSMTEST |
Clip State Machine Test and Update |
DEPBAR |
Dependency Barrier |
GETLMEMBASE |
Get Local Memory Base Address |
LEPC |
Load Effective PC |
NOP |
No Operation |
PMTRIG |
Performance Monitor Trigger |
R2B |
Move Register to Barrier |
S2R |
Move Special Register to Register |
SETCTAID |
Set CTA ID |
SETLMEMBASE |
Set Local Memory Base Address |
VOTE |
Vote Across SIMD Thread Group |
VOTE_VTG |
Clip State Machine Test and Update |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FADD32I |
FP32 Add |
FCHK |
Floating-point Range Check |
FFMA32I |
FP32 Fused Multiply and Add |
FFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMUL |
FP32 Multiply |
FMUL32I |
FP32 Multiply |
FSEL |
Floating Point Select |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Swizzle Add |
MUFU |
FP32 Multi Function Operation |
HADD2 |
FP16 Add |
HADD2_32I |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HFMA2_32I |
FP16 Fused Mutiply Add |
HMMA |
Matrix Multiply and Accumulate |
HMUL2 |
FP16 Multiply |
HMUL2_32I |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMUL |
FP64 Multiply |
DSETP |
FP64 Compare And Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
BMMA |
Bit Matrix Multiply and Accumulate |
BMSK |
Bitfield Mask |
BREV |
Bit Reverse |
FLO |
Find Leading One |
IABS |
Integer Absolute Value |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
IADD32I |
Integer Addition |
IDP |
Integer Dot Product and Accumulate |
IDP4A |
Integer Dot Product and Accumulate |
IMAD |
Integer Multiply And Add |
IMMA |
Integer Matrix Multiply and Accumulate |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
IMUL32I |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISCADD32I |
Scaled Integer Addition |
ISETP |
Integer Compare And Set Predicate |
LEA |
LOAD Effective Address |
LOP |
Logic Operation |
LOP3 |
Logic Operation |
LOP32I |
Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
VABSDIFF |
Absolute Difference |
VABSDIFF4 |
Absolute Difference |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
I2IP |
Integer To Integer Conversion and Packing |
FRND |
Round To Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
MOV32I |
Move |
MOVM |
Move Matrix with Transposition or Expansion |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SGXT |
Sign Extend |
SHFL |
Warp Wide Register Shuffle |
Predicate Instructions
| Opcode | Description |
|---|---|
PLOP3 |
Predicate Logic Operation |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate Register |
Load/Store Instructions
| Opcode | Description |
|---|---|
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDL |
Load within Local Memory Window |
LDS |
Load within Shared Memory Window |
LDSM |
Load Matrix from Shared Memory with Element Size Expansion |
ST |
Store to Generic Memory |
STG |
Store to Global Memory |
STL |
Store within Local or Shared Window |
STS |
Store within Local or Shared Window |
MATCH |
Match Register Values Across Thread Group |
QSPC |
Query Space |
ATOM |
Atomic Operation on Generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
ATOMG |
Atomic Operation on Global Memory |
RED |
Reduction Operation on Generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
ERRBAR |
Error Barrier |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Uniform Datapath Instructions
| Opcode | Description |
|---|---|
R2UR |
Move from Vector Register to a Uniform Register |
S2UR |
Move Special Register to Uniform Register |
UBMSK |
Uniform Bitfield Mask |
UBREV |
Uniform Bit Reverse |
UCLEA |
Load Effective Address for a Constant |
UFLO |
Uniform Find Leading One |
UIADD3 |
Uniform Integer Addition |
UIADD3.64 |
Uniform Integer Addition |
UIMAD |
Uniform Integer Multiplication |
UISETP |
Integer Compare and Set Uniform Predicate |
ULDC |
Load from Constant Memory into a Uniform Register |
ULEA |
Uniform Load Effective Address |
ULOP |
Logic Operation |
ULOP3 |
Logic Operation |
ULOP32I |
Logic Operation |
UMOV |
Uniform Move |
UP2UR |
Uniform Predicate to Uniform Register |
UPLOP3 |
Uniform Predicate Logic Operation |
UPOPC |
Uniform Population Count |
UPRMT |
Uniform Byte Permute |
UPSETP |
Uniform Predicate Logic Operation |
UR2UP |
Uniform Register to Uniform Predicate |
USEL |
Uniform Select |
USGXT |
Uniform Sign Extend |
USHF |
Uniform Funnel Shift |
USHL |
Uniform Left Shift |
USHR |
Uniform Right Shift |
VOTEU |
Voting across SIMD Thread Group with Results in Uniform Destination |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TMML |
Texture MipMap Level |
TXD |
Texture Fetch With Derivatives |
TXQ |
Texture Query |
Surface Instructions
| Opcode | Description |
|---|---|
SUATOM |
Surface Reduction |
SULD |
Surface Load |
SURED |
Atomic Reduction on surface memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
BMOV |
Move Convergence Barrier State |
BPT |
BreakPoint/Trap |
BRA |
Relative Branch |
BREAK |
Break out of the Specified Convergence Barrier |
BRX |
Relative Branch Indirect |
BRXU |
Relative Branch with Uniform Register Based Offset |
BSSY |
Barrier Set Convergence Synchronization Point |
BSYNC |
Synchronize Threads on a Convergence Barrier |
CALL |
Call Function |
EXIT |
Exit Program |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
JMXU |
Absolute Jump with Uniform Register Based Offset |
KILL |
Kill Thread |
NANOSLEEP |
Suspend Execution |
RET |
Return From Subroutine |
RPCMOV |
PC Register Move |
RTT |
Return From Trap |
WARPSYNC |
Synchronize Threads in Warp |
YIELD |
Yield Control |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
CS2R |
Move Special Register to Register |
CSMTEST |
Clip State Machine Test and Update |
DEPBAR |
Dependency Barrier |
GETLMEMBASE |
Get Local Memory Base Address |
LEPC |
Load Effective PC |
NOP |
No Operation |
PMTRIG |
Performance Monitor Trigger |
R2B |
Move Register to Barrier |
S2R |
Move Special Register to Register |
SETCTAID |
Set CTA ID |
SETLMEMBASE |
Set Local Memory Base Address |
VOTE |
Vote Across SIMD Thread Group |
VOTE_VTG |
Clip State Machine Test and Update |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FADD32I |
FP32 Add |
FCHK |
Floating-point Range Check |
FFMA32I |
FP32 Fused Multiply and Add |
FFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMUL |
FP32 Multiply |
FMUL32I |
FP32 Multiply |
FSEL |
Floating Point Select |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Swizzle Add |
MUFU |
FP32 Multi Function Operation |
HADD2 |
FP16 Add |
HADD2_32I |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HFMA2_32I |
FP16 Fused Mutiply Add |
HMMA |
Matrix Multiply and Accumulate |
HMUL2 |
FP16 Multiply |
HMUL2_32I |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMUL |
FP64 Multiply |
DSETP |
FP64 Compare And Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
BMMA |
Bit Matrix Multiply and Accumulate |
BMSK |
Bitfield Mask |
BREV |
Bit Reverse |
FLO |
Find Leading One |
IABS |
Integer Absolute Value |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
IADD32I |
Integer Addition |
IDP |
Integer Dot Product and Accumulate |
IDP4A |
Integer Dot Product and Accumulate |
IMAD |
Integer Multiply And Add |
IMMA |
Integer Matrix Multiply and Accumulate |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
IMUL32I |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISCADD32I |
Scaled Integer Addition |
ISETP |
Integer Compare And Set Predicate |
LEA |
LOAD Effective Address |
LOP |
Logic Operation |
LOP3 |
Logic Operation |
LOP32I |
Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
VABSDIFF |
Absolute Difference |
VABSDIFF4 |
Absolute Difference |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
I2IP |
Integer To Integer Conversion and Packing |
FRND |
Round To Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
MOV32I |
Move |
MOVM |
Move Matrix with Transposition or Expansion |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SGXT |
Sign Extend |
SHFL |
Warp Wide Register Shuffle |
Predicate Instructions
| Opcode | Description |
|---|---|
PLOP3 |
Predicate Logic Operation |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate Register |
Load/Store Instructions
| Opcode | Description |
|---|---|
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDL |
Load within Local Memory Window |
LDS |
Load within Shared Memory Window |
LDSM |
Load Matrix from Shared Memory with Element Size Expansion |
ST |
Store to Generic Memory |
STG |
Store to Global Memory |
STL |
Store to Local Memory |
STS |
Store to Shared Memory |
MATCH |
Match Register Values Across Thread Group |
QSPC |
Query Space |
ATOM |
Atomic Operation on Generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
ATOMG |
Atomic Operation on Global Memory |
RED |
Reduction Operation on Generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
ERRBAR |
Error Barrier |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Uniform Datapath Instructions
| Opcode | Description |
|---|---|
R2UR |
Move from Vector Register to a Uniform Register |
S2UR |
Move Special Register to Uniform Register |
UBMSK |
Uniform Bitfield Mask |
UBREV |
Uniform Bit Reverse |
UCLEA |
Load Effective Address for a Constant |
UFLO |
Uniform Find Leading One |
UIADD3 |
Uniform Integer Addition |
UIADD3.64 |
Uniform Integer Addition |
UIMAD |
Uniform Integer Multiplication |
UISETP |
Integer Compare and Set Uniform Predicate |
ULDC |
Load from Constant Memory into a Uniform Register |
ULEA |
Uniform Load Effective Address |
ULOP |
Logic Operation |
ULOP3 |
Logic Operation |
ULOP32I |
Logic Operation |
UMOV |
Uniform Move |
UP2UR |
Uniform Predicate to Uniform Register |
UPLOP3 |
Uniform Predicate Logic Operation |
UPOPC |
Uniform Population Count |
UPRMT |
Uniform Byte Permute |
UPSETP |
Uniform Predicate Logic Operation |
UR2UP |
Uniform Register to Uniform Predicate |
USEL |
Uniform Select |
USGXT |
Uniform Sign Extend |
USHF |
Uniform Funnel Shift |
USHL |
Uniform Left Shift |
USHR |
Uniform Right Shift |
VOTEU |
Voting across SIMD Thread Group with Results in Uniform Destination |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TMML |
Texture MipMap Level |
TXD |
Texture Fetch With Derivatives |
TXQ |
Texture Query |
Surface Instructions
| Opcode | Description |
|---|---|
SUATOM |
Atomic Op on Surface Memory |
SULD |
Surface Load |
SURED |
Reduction Op on Surface Memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
BMOV |
Move Convergence Barrier State |
BPT |
BreakPoint/Trap |
BRA |
Relative Branch |
BREAK |
Break out of the Specified Convergence Barrier |
BRX |
Relative Branch Indirect |
BRXU |
Relative Branch with Uniform Register Based Offset |
BSSY |
Barrier Set Convergence Synchronization Point |
BSYNC |
Synchronize Threads on a Convergence Barrier |
CALL |
Call Function |
EXIT |
Exit Program |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
JMXU |
Absolute Jump with Uniform Register Based Offset |
KILL |
Kill Thread |
NANOSLEEP |
Suspend Execution |
RET |
Return From Subroutine |
RPCMOV |
PC Register Move |
RTT |
Return From Trap |
WARPSYNC |
Synchronize Threads in Warp |
YIELD |
Yield Control |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
CS2R |
Move Special Register to Register |
DEPBAR |
Dependency Barrier |
GETLMEMBASE |
Get Local Memory Base Address |
LEPC |
Load Effective PC |
NOP |
No Operation |
PMTRIG |
Performance Monitor Trigger |
R2B |
Move Register to Barrier |
S2R |
Move Special Register to Register |
SETCTAID |
Set CTA ID |
SETLMEMBASE |
Set Local Memory Base Address |
VOTE |
Vote Across SIMD Thread Group |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FADD32I |
FP32 Add |
FCHK |
Floating-point Range Check |
FFMA32I |
FP32 Fused Multiply and Add |
FFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMUL |
FP32 Multiply |
FMUL32I |
FP32 Multiply |
FSEL |
Floating Point Select |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Swizzle Add |
MUFU |
FP32 Multi Function Operation |
HADD2 |
FP16 Add |
HADD2_32I |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HFMA2_32I |
FP16 Fused Mutiply Add |
HMMA |
Matrix Multiply and Accumulate |
HMNMX2 |
FP16 Minimum / Maximum |
HMUL2 |
FP16 Multiply |
HMUL2_32I |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMMA |
Matrix Multiply and Accumulate |
DMUL |
FP64 Multiply |
DSETP |
FP64 Compare And Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
BMMA |
Bit Matrix Multiply and Accumulate |
BMSK |
Bitfield Mask |
BREV |
Bit Reverse |
FLO |
Find Leading One |
IABS |
Integer Absolute Value |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
IADD32I |
Integer Addition |
IDP |
Integer Dot Product and Accumulate |
IDP4A |
Integer Dot Product and Accumulate |
IMAD |
Integer Multiply And Add |
IMMA |
Integer Matrix Multiply and Accumulate |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
IMUL32I |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISCADD32I |
Scaled Integer Addition |
ISETP |
Integer Compare And Set Predicate |
LEA |
LOAD Effective Address |
LOP |
Logic Operation |
LOP3 |
Logic Operation |
LOP32I |
Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
VABSDIFF |
Absolute Difference |
VABSDIFF4 |
Absolute Difference |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
I2IP |
Integer To Integer Conversion and Packing |
I2FP |
Integer to FP32 Convert and Pack |
F2IP |
FP32 Down-Convert to Integer and Pack |
FRND |
Round To Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
MOV32I |
Move |
MOVM |
Move Matrix with Transposition or Expansion |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SGXT |
Sign Extend |
SHFL |
Warp Wide Register Shuffle |
Predicate Instructions
| Opcode | Description |
|---|---|
PLOP3 |
Predicate Logic Operation |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate Register |
Load/Store Instructions
| Opcode | Description |
|---|---|
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDGDEPBAR |
Global Load Dependency Barrier |
LDGSTS |
Asynchronous Global to Shared Memcopy |
LDL |
Load within Local Memory Window |
LDS |
Load within Shared Memory Window |
LDSM |
Load Matrix from Shared Memory with Element Size Expansion |
ST |
Store to Generic Memory |
STG |
Store to Global Memory |
STL |
Store to Local Memory |
STS |
Store to Shared Memory |
MATCH |
Match Register Values Across Thread Group |
QSPC |
Query Space |
ATOM |
Atomic Operation on Generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
ATOMG |
Atomic Operation on Global Memory |
RED |
Reduction Operation on Generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
ERRBAR |
Error Barrier |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Uniform Datapath Instructions
| Opcode | Description |
|---|---|
R2UR |
Move from Vector Register to a Uniform Register |
REDUX |
Reduction of a Vector Register into a Uniform Register |
S2UR |
Move Special Register to Uniform Register |
UBMSK |
Uniform Bitfield Mask |
UBREV |
Uniform Bit Reverse |
UCLEA |
Load Effective Address for a Constant |
UF2FP |
Uniform FP32 Down-convert and Pack |
UFLO |
Uniform Find Leading One |
UIADD3 |
Uniform Integer Addition |
UIADD3.64 |
Uniform Integer Addition |
UIMAD |
Uniform Integer Multiplication |
UISETP |
Integer Compare and Set Uniform Predicate |
ULDC |
Load from Constant Memory into a Uniform Register |
ULEA |
Uniform Load Effective Address |
ULOP |
Logic Operation |
ULOP3 |
Logic Operation |
ULOP32I |
Logic Operation |
UMOV |
Uniform Move |
UP2UR |
Uniform Predicate to Uniform Register |
UPLOP3 |
Uniform Predicate Logic Operation |
UPOPC |
Uniform Population Count |
UPRMT |
Uniform Byte Permute |
UPSETP |
Uniform Predicate Logic Operation |
UR2UP |
Uniform Register to Uniform Predicate |
USEL |
Uniform Select |
USGXT |
Uniform Sign Extend |
USHF |
Uniform Funnel Shift |
USHL |
Uniform Left Shift |
USHR |
Uniform Right Shift |
VOTEU |
Voting across SIMD Thread Group with Results in Uniform Destination |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TMML |
Texture MipMap Level |
TXD |
Texture Fetch With Derivatives |
TXQ |
Texture Query |
Surface Instructions
| Opcode | Description |
|---|---|
SUATOM |
Atomic Op on Surface Memory |
SULD |
Surface Load |
SURED |
Reduction Op on Surface Memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
BMOV |
Move Convergence Barrier State |
BPT |
BreakPoint/Trap |
BRA |
Relative Branch |
BREAK |
Break out of the Specified Convergence Barrier |
BRX |
Relative Branch Indirect |
BRXU |
Relative Branch with Uniform Register Based Offset |
BSSY |
Barrier Set Convergence Synchronization Point |
BSYNC |
Synchronize Threads on a Convergence Barrier |
CALL |
Call Function |
EXIT |
Exit Program |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
JMXU |
Absolute Jump with Uniform Register Based Offset |
KILL |
Kill Thread |
NANOSLEEP |
Suspend Execution |
RET |
Return From Subroutine |
RPCMOV |
PC Register Move |
WARPSYNC |
Synchronize Threads in Warp |
YIELD |
Yield Control |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
CS2R |
Move Special Register to Register |
DEPBAR |
Dependency Barrier |
GETLMEMBASE |
Get Local Memory Base Address |
LEPC |
Load Effective PC |
NOP |
No Operation |
PMTRIG |
Performance Monitor Trigger |
S2R |
Move Special Register to Register |
SETCTAID |
Set CTA ID |
SETLMEMBASE |
Set Local Memory Base Address |
VOTE |
Vote Across SIMD Thread Group |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FADD32I |
FP32 Add |
FCHK |
Floating-point Range Check |
FFMA32I |
FP32 Fused Multiply and Add |
FFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMUL |
FP32 Multiply |
FMUL32I |
FP32 Multiply |
FSEL |
Floating Point Select |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Swizzle Add |
MUFU |
FP32 Multi Function Operation |
HADD2 |
FP16 Add |
HADD2_32I |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HFMA2_32I |
FP16 Fused Mutiply Add |
HMMA |
Matrix Multiply and Accumulate |
HMNMX2 |
FP16 Minimum / Maximum |
HMUL2 |
FP16 Multiply |
HMUL2_32I |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMMA |
Matrix Multiply and Accumulate |
DMUL |
FP64 Multiply |
DSETP |
FP64 Compare And Set Predicate |
Integer Instructions
| Opcode | Description |
|---|---|
BMMA |
Bit Matrix Multiply and Accumulate |
BMSK |
Bitfield Mask |
BREV |
Bit Reverse |
FLO |
Find Leading One |
IABS |
Integer Absolute Value |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
IADD32I |
Integer Addition |
IDP |
Integer Dot Product and Accumulate |
IDP4A |
Integer Dot Product and Accumulate |
IMAD |
Integer Multiply And Add |
IMMA |
Integer Matrix Multiply and Accumulate |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
IMUL32I |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISCADD32I |
Scaled Integer Addition |
ISETP |
Integer Compare And Set Predicate |
LEA |
LOAD Effective Address |
LOP |
Logic Operation |
LOP3 |
Logic Operation |
LOP32I |
Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
VABSDIFF |
Absolute Difference |
VABSDIFF4 |
Absolute Difference |
VHMNMX |
SIMD FP16 3-Input Minimum / Maximum |
VIADD |
SIMD Integer Addition |
VIADDMNMX |
SIMD Integer Addition and Fused Min/Max Comparison |
VIMNMX |
SIMD Integer Minimum / Maximum |
VIMNMX3 |
SIMD Integer 3-Input Minimum / Maximum |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
I2IP |
Integer To Integer Conversion and Packing |
I2FP |
Integer to FP32 Convert and Pack |
F2IP |
FP32 Down-Convert to Integer and Pack |
FRND |
Round To Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
MOV32I |
Move |
MOVM |
Move Matrix with Transposition or Expansion |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SGXT |
Sign Extend |
SHFL |
Warp Wide Register Shuffle |
Predicate Instructions
| Opcode | Description |
|---|---|
PLOP3 |
Predicate Logic Operation |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate Register |
Load/Store Instructions
| Opcode | Description |
|---|---|
FENCE |
Memory Visibility Guarantee for Shared or Global Memory |
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDGDEPBAR |
Global Load Dependency Barrier |
LDGMC |
Reducing Load |
LDGSTS |
Asynchronous Global to Shared Memcopy |
LDL |
Load within Local Memory Window |
LDS |
Load within Shared Memory Window |
LDSM |
Load Matrix from Shared Memory with Element Size Expansion |
STSM |
Store Matrix to Shared Memory |
ST |
Store to Generic Memory |
STG |
Store to Global Memory |
STL |
Store to Local Memory |
STS |
Store to Shared Memory |
STAS |
Asynchronous Store to Distributed Shared Memory With Explicit Synchronization |
SYNCS |
Sync Unit |
MATCH |
Match Register Values Across Thread Group |
QSPC |
Query Space |
ATOM |
Atomic Operation on Generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
ATOMG |
Atomic Operation on Global Memory |
REDAS |
Asynchronous Reduction on Distributed Shared Memory With Explicit Synchronization |
REDG |
Reduction Operation on Generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
ERRBAR |
Error Barrier |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Uniform Datapath Instructions
| Opcode | Description |
|---|---|
R2UR |
Move from Vector Register to a Uniform Register |
REDUX |
Reduction of a Vector Register into a Uniform Register |
S2UR |
Move Special Register to Uniform Register |
UBMSK |
Uniform Bitfield Mask |
UBREV |
Uniform Bit Reverse |
UCGABAR_ARV |
CGA Barrier Synchronization |
UCGABAR_WAIT |
CGA Barrier Synchronization |
UCLEA |
Load Effective Address for a Constant |
UF2FP |
Uniform FP32 Down-convert and Pack |
UFLO |
Uniform Find Leading One |
UIADD3 |
Uniform Integer Addition |
UIADD3.64 |
Uniform Integer Addition |
UIMAD |
Uniform Integer Multiplication |
UISETP |
Integer Compare and Set Uniform Predicate |
ULDC |
Load from Constant Memory into a Uniform Register |
ULEA |
Uniform Load Effective Address |
ULEPC |
Uniform Load Effective PC |
ULOP |
Logic Operation |
ULOP3 |
Logic Operation |
ULOP32I |
Logic Operation |
UMOV |
Uniform Move |
UP2UR |
Uniform Predicate to Uniform Register |
UPLOP3 |
Uniform Predicate Logic Operation |
UPOPC |
Uniform Population Count |
UPRMT |
Uniform Byte Permute |
UPSETP |
Uniform Predicate Logic Operation |
UR2UP |
Uniform Register to Uniform Predicate |
USEL |
Uniform Select |
USETMAXREG |
Release, Deallocate and Allocate Registers |
USGXT |
Uniform Sign Extend |
USHF |
Uniform Funnel Shift |
USHL |
Uniform Left Shift |
USHR |
Uniform Right Shift |
VOTEU |
Voting across SIMD Thread Group with Results in Uniform Destination |
Warpgroup Instructions
| Opcode | Description |
|---|---|
BGMMA |
Bit Matrix Multiply and Accumulate Across Warps |
HGMMA |
Matrix Multiply and Accumulate Across a Warpgroup |
IGMMA |
Integer Matrix Multiply and Accumulate Across a Warpgroup |
QGMMA |
FP8 Matrix Multiply and Accumulate Across a Warpgroup |
WARPGROUP |
Warpgroup Synchronization |
WARPGROUPSET |
Set Warpgroup Counters |
Tensor Memory Access Instructions
| Opcode | Description |
|---|---|
UBLKCP |
Bulk Data Copy |
UBLKPF |
Bulk Data Prefetch |
UBLKRED |
Bulk Data Copy from Shared Memory with Reduction |
UTMACCTL |
TMA Cache Control |
UTMACMDFLUSH |
TMA Command Flush |
UTMALDG |
Tensor Load from Global to Shared Memory |
UTMAPF |
Tensor Prefetch |
UTMAREDG |
Tensor Store from Shared to Global Memory with Reduction |
UTMASTG |
Tensor Store from Shared to Global Memory |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TMML |
Texture MipMap Level |
TXD |
Texture Fetch With Derivatives |
TXQ |
Texture Query |
Surface Instructions
| Opcode | Description |
|---|---|
SUATOM |
Atomic Op on Surface Memory |
SULD |
Surface Load |
SURED |
Reduction Op on Surface Memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
ACQBULK |
Wait for Bulk Release Status Warp State |
BMOV |
Move Convergence Barrier State |
BPT |
BreakPoint/Trap |
BRA |
Relative Branch |
BREAK |
Break out of the Specified Convergence Barrier |
BRX |
Relative Branch Indirect |
BRXU |
Relative Branch with Uniform Register Based Offset |
BSSY |
Barrier Set Convergence Synchronization Point |
BSYNC |
Synchronize Threads on a Convergence Barrier |
CALL |
Call Function |
CGAERRBAR |
CGA Error Barrier |
ELECT |
Elect a Leader Thread |
ENDCOLLECTIVE |
Reset the MCOLLECTIVE mask |
EXIT |
Exit Program |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
JMXU |
Absolute Jump with Uniform Register Based Offset |
KILL |
Kill Thread |
NANOSLEEP |
Suspend Execution |
PREEXIT |
Dependent Task Launch Hint |
RET |
Return From Subroutine |
RPCMOV |
PC Register Move |
WARPSYNC |
Synchronize Threads in Warp |
YIELD |
Yield Control |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
CS2R |
Move Special Register to Register |
DEPBAR |
Dependency Barrier |
GETLMEMBASE |
Get Local Memory Base Address |
LEPC |
Load Effective PC |
NOP |
No Operation |
PMTRIG |
Performance Monitor Trigger |
S2R |
Move Special Register to Register |
SETCTAID |
Set CTA ID |
SETLMEMBASE |
Set Local Memory Base Address |
VOTE |
Vote Across SIMT Thread Group |
Floating Point Instructions
| Opcode | Description |
|---|---|
FADD |
FP32 Add |
FADD2 |
FP32 Add |
FADD32I |
FP32 Add |
FCHK |
Floating-point Range Check |
FFMA32I |
FP32 Fused Multiply and Add |
FFMA |
FP32 Fused Multiply and Add |
FFMA2 |
FP32 Fused Multiply and Add |
FHADD |
FP32 Addition |
FHFMA |
FP32 Fused Multiply and Add |
FMNMX |
FP32 Minimum/Maximum |
FMNMX3 |
3-Input Floating-point Minimum / Maximum |
FMUL |
FP32 Multiply |
FMUL2 |
FP32 Multiply |
FMUL32I |
FP32 Multiply |
FSEL |
Floating Point Select |
FSET |
FP32 Compare And Set |
FSETP |
FP32 Compare And Set Predicate |
FSWZADD |
FP32 Swizzle Add |
MUFU |
FP32 Multi Function Operation |
HADD2 |
FP16 Add |
HADD2_32I |
FP16 Add |
HFMA2 |
FP16 Fused Mutiply Add |
HFMA2_32I |
FP16 Fused Mutiply Add |
HMMA |
Matrix Multiply and Accumulate |
HMNMX2 |
FP16 Minimum / Maximum |
HMUL2 |
FP16 Multiply |
HMUL2_32I |
FP16 Multiply |
HSET2 |
FP16 Compare And Set |
HSETP2 |
FP16 Compare And Set Predicate |
DADD |
FP64 Add |
DFMA |
FP64 Fused Mutiply Add |
DMMA |
Matrix Multiply and Accumulate |
DMUL |
FP64 Multiply |
DSETP |
FP64 Compare And Set Predicate |
OMMA |
FP4 Matrix Multiply and Accumulate Across a Warp |
QMMA |
FP8 Matrix Multiply and Accumulate Across a Warp |
Integer Instructions
| Opcode | Description |
|---|---|
BMSK |
Bitfield Mask |
BREV |
Bit Reverse |
FLO |
Find Leading One |
IABS |
Integer Absolute Value |
IADD |
Integer Addition |
IADD3 |
3-input Integer Addition |
IADD32I |
Integer Addition |
IDP |
Integer Dot Product and Accumulate |
IDP4A |
Integer Dot Product and Accumulate |
IMAD |
Integer Multiply And Add |
IMMA |
Integer Matrix Multiply and Accumulate |
IMNMX |
Integer Minimum/Maximum |
IMUL |
Integer Multiply |
IMUL32I |
Integer Multiply |
ISCADD |
Scaled Integer Addition |
ISCADD32I |
Scaled Integer Addition |
ISETP |
Integer Compare And Set Predicate |
LEA |
LOAD Effective Address |
LOP |
Logic Operation |
LOP3 |
Logic Operation |
LOP32I |
Logic Operation |
POPC |
Population count |
SHF |
Funnel Shift |
SHL |
Shift Left |
SHR |
Shift Right |
VABSDIFF |
Absolute Difference |
VABSDIFF4 |
Absolute Difference |
VHMNMX |
SIMD FP16 3-Input Minimum / Maximum |
VIADD |
SIMD Integer Addition |
VIADDMNMX |
SIMD Integer Addition and Fused Min/Max Comparison |
VIMNMX |
SIMD Integer Minimum / Maximum |
VIMNMX3 |
SIMD Integer 3-Input Minimum / Maximum |
Conversion Instructions
| Opcode | Description |
|---|---|
F2F |
Floating Point To Floating Point Conversion |
F2I |
Floating Point To Integer Conversion |
I2F |
Integer To Floating Point Conversion |
I2I |
Integer To Integer Conversion |
I2IP |
Integer To Integer Conversion and Packing |
I2FP |
Integer to FP32 Convert and Pack |
F2IP |
FP32 Down-Convert to Integer and Pack |
FRND |
Round To Integer |
Movement Instructions
| Opcode | Description |
|---|---|
MOV |
Move |
MOV32I |
Move |
MOVM |
Move Matrix with Transposition or Expansion |
PRMT |
Permute Register Pair |
SEL |
Select Source with Predicate |
SGXT |
Sign Extend |
SHFL |
Warp Wide Register Shuffle |
Predicate Instructions
| Opcode | Description |
|---|---|
PLOP3 |
Predicate Logic Operation |
PSETP |
Combine Predicates and Set Predicate |
P2R |
Move Predicate Register To Register |
R2P |
Move Register To Predicate Register |
Load/Store Instructions
| Opcode | Description |
|---|---|
FENCE |
Memory Visibility Guarantee for Shared or Global Memory |
LD |
Load from generic Memory |
LDC |
Load Constant |
LDG |
Load from Global Memory |
LDGDEPBAR |
Global Load Dependency Barrier |
LDGMC |
Reducing Load |
LDGSTS |
Asynchronous Global to Shared Memcopy |
LDL |
Load within Local Memory Window |
LDS |
Load within Shared Memory Window |
LDSM |
Load Matrix from Shared Memory with Element Size Expansion |
STSM |
Store Matrix to Shared Memory |
ST |
Store to Generic Memory |
STG |
Store to Global Memory |
STL |
Store to Local Memory |
STS |
Store to Shared Memory |
STAS |
Asynchronous Store to Distributed Shared Memory With Explicit Synchronization |
SYNCS |
Sync Unit |
MATCH |
Match Register Values Across Thread Group |
QSPC |
Query Space |
ATOM |
Atomic Operation on Generic Memory |
ATOMS |
Atomic Operation on Shared Memory |
ATOMG |
Atomic Operation on Global Memory |
REDAS |
Asynchronous Reduction on Distributed Shared Memory With Explicit Synchronization |
REDG |
Reduction Operation on Generic Memory |
CCTL |
Cache Control |
CCTLL |
Cache Control |
ERRBAR |
Error Barrier |
MEMBAR |
Memory Barrier |
CCTLT |
Texture Cache Control |
Uniform Datapath Instructions
| Opcode | Description |
|---|---|
CREDUX |
Coupled Reduction of a Vector Register into a Uniform Register |
CS2UR |
Load a Value from Constant Memory into a Uniform Register |
LDCU |
Load a Value from Constant Memory into a Uniform Register |
R2UR |
Move from Vector Register to a Uniform Register |
REDUX |
Reduction of a Vector Register into a Uniform Register |
S2UR |
Move Special Register to Uniform Register |
UBMSK |
Uniform Bitfield Mask |
UBREV |
Uniform Bit Reverse |
UCGABAR_ARV |
CGA Barrier Synchronization |
UCGABAR_WAIT |
CGA Barrier Synchronization |
UCLEA |
Load Effective Address for a Constant |
UFADD |
Uniform Uniform FP32 Addition |
UF2F |
Uniform Float-to-Float Conversion |
UF2FP |
Uniform FP32 Down-convert and Pack |
UF2I |
Uniform Float-to-Integer Conversion |
UF2IP |
Uniform FP32 Down-Convert to Integer and Pack |
UFFMA |
Uniform FP32 Fused Multiply-Add |
UFLO |
Uniform Find Leading One |
UFMNMX |
Uniform Floating-point Minimum / Maximum |
UFMUL |
Uniform FP32 Multiply |
UFRND |
Uniform Round to Integer |
UFSEL |
Uniform Floating-Point Select |
UFSET |
Uniform Floating-Point Compare and Set |
UFSETP |
Uniform Floating-Point Compare and Set Predicate |
UI2F |
Uniform Integer to Float conversion |
UI2FP |
Uniform Integer to FP32 Convert and Pack |
UI2I |
Uniform Saturating Integer-to-Integer Conversion |
UI2IP |
Uniform Dual Saturating Integer-to-Integer Conversion and Packing |
UIABS |
Uniform Integer Absolute Value |
UIMNMX |
Uniform Integer Minimum / Maximum |
UIADD3 |
Uniform Integer Addition |
UIADD3.64 |
Uniform Integer Addition |
UIMAD |
Uniform Integer Multiplication |
UISETP |
Uniform Integer Compare and Set Uniform Predicate |
ULEA |
Uniform Load Effective Address |
ULEPC |
Uniform Load Effective PC |
ULOP |
Uniform Logic Operation |
ULOP3 |
Uniform Logic Operation |
ULOP32I |
Uniform Logic Operation |
UMOV |
Uniform Move |
UP2UR |
Uniform Predicate to Uniform Register |
UPLOP3 |
Uniform Predicate Logic Operation |
UPOPC |
Uniform Population Count |
UPRMT |
Uniform Byte Permute |
UPSETP |
Uniform Predicate Logic Operation |
UR2UP |
Uniform Register to Uniform Predicate |
USEL |
Uniform Select |
USETMAXREG |
Release, Deallocate and Allocate Registers |
USGXT |
Uniform Sign Extend |
USHF |
Uniform Funnel Shift |
USHL |
Uniform Left Shift |
USHR |
Uniform Right Shift |
UGETNEXTWORKID |
Uniform Get Next Work ID |
UMEMSETS |
Initialize Shared Memory |
UREDGR |
Uniform Reduction on Global Memory with Release |
USTGR |
Uniform Store to Global Memory with Release |
UVIADD |
Uniform SIMD Integer Addition |
UVIMNMX |
Uniform SIMD Integer Minimum / Maximum |
UVIRTCOUNT |
Virtual Resource Management |
VOTEU |
Voting across SIMD Thread Group with Results in Uniform Destination |
Tensor Memory Access Instructions
| Opcode | Description |
|---|---|
UBLKCP |
Bulk Data Copy |
UBLKPF |
Bulk Data Prefetch |
UBLKRED |
Bulk Data Copy from Shared Memory with Reduction |
UTMACCTL |
TMA Cache Control |
UTMACMDFLUSH |
TMA Command Flush |
UTMALDG |
Tensor Load from Global to Shared Memory |
UTMAPF |
Tensor Prefetch |
UTMAREDG |
Tensor Store from Shared to Global Memory with Reduction |
UTMASTG |
Tensor Store from Shared to Global Memory |
Tensor Core Memory Instructions
| Opcode | Description |
|---|---|
LDT |
Load Matrix from Tensor Memory to Register File |
LDTM |
Load Matrix from Tensor Memory to Register File |
STT |
Store Matrix to Tensor Memory from Register File |
STTM |
Store Matrix to Tensor Memory from Register File |
UTCATOMSWS |
Perform Atomic operation on SW State Register |
UTCBAR |
Tensor Core Barrier |
UTCCP |
Asynchonous data copy from Shared Memory to Tensor Memory |
UTCHMMA |
Uniform Matrix Multiply and Accumulate |
UTCIMMA |
Uniform Matrix Multiply and Accumulate |
UTCOMMA |
Uniform Matrix Multiply and Accumulate |
UTCQMMA |
Uniform Matrix Multiply and Accumulate |
UTCSHIFT |
Shift elements in Tensor Memory |
Texture Instructions
| Opcode | Description |
|---|---|
TEX |
Texture Fetch |
TLD |
Texture Load |
TLD4 |
Texture Load 4 |
TMML |
Texture MipMap Level |
TXD |
Texture Fetch With Derivatives |
TXQ |
Texture Query |
Surface Instructions
| Opcode | Description |
|---|---|
SUATOM |
Atomic Op on Surface Memory |
SULD |
Surface Load |
SURED |
Reduction Op on Surface Memory |
SUST |
Surface Store |
Control Instructions
| Opcode | Description |
|---|---|
ACQBULK |
Wait for Bulk Release Status Warp State |
ACQSHMINIT |
Wait for Shared Memory Initialization Release Status Warp State |
BMOV |
Move Convergence Barrier State |
BPT |
BreakPoint/Trap |
BRA |
Relative Branch |
BREAK |
Break out of the Specified Convergence Barrier |
BRX |
Relative Branch Indirect |
BRXU |
Relative Branch with Uniform Register Based Offset |
BSSY |
Barrier Set Convergence Synchronization Point |
BSYNC |
Synchronize Threads on a Convergence Barrier |
CALL |
Call Function |
CGAERRBAR |
CGA Error Barrier |
ELECT |
Elect a Leader Thread |
ENDCOLLECTIVE |
Reset the MCOLLECTIVE mask |
EXIT |
Exit Program |
JMP |
Absolute Jump |
JMX |
Absolute Jump Indirect |
JMXU |
Absolute Jump with Uniform Register Based Offset |
KILL |
Kill Thread |
NANOSLEEP |
Suspend Execution |
PREEXIT |
Dependent Task Launch Hint |
RET |
Return From Subroutine |
RPCMOV |
PC Register Move |
WARPSYNC |
Synchronize Threads in Warp |
YIELD |
Yield Control |
Miscellaneous Instructions
| Opcode | Description |
|---|---|
B2R |
Move Barrier To Register |
BAR |
Barrier Synchronization |
CS2R |
Move Special Register to Register |
DEPBAR |
Dependency Barrier |
GETLMEMBASE |
Get Local Memory Base Address |
LEPC |
Load Effective PC |
NOP |
No Operation |
PMTRIG |
Performance Monitor Trigger |
S2R |
Move Special Register to Register |
SETCTAID |
Set CTA ID |
SETLMEMBASE |
Set Local Memory Base Address |
VOTE |
Vote Across SIMT Thread Group |