Module x86_64
x86_64 intrinsics
Modules
- abm Advanced Bit Manipulation (ABM) instructions
- adx
- amx
- avx Advanced Vector Extensions (AVX)
- avx512bw
- avx512f
- avx512fp16
- bmi Bit Manipulation Instruction (BMI) Set 1.0.
- bmi2 Bit Manipulation Instruction (BMI) Set 2.0.
- bswap Byte swap intrinsics.
- bt
- cmpxchg16b
- fxsr FXSR floating-point context fast save and restore.
- macros Utility macros.
- movrs Read-shared Move instructions
- rdrand RDRAND and RDSEED instructions for returning random numbers from an Intel on-chip hardware random number generator which has been seeded by an on-chip entropy source.
-
sse
x86_64Streaming SIMD Extensions (SSE) -
sse2
x86_64's Streaming SIMD Extensions 2 (SSE2) -
sse41
i686's Streaming SIMD Extensions 4.1 (SSE4.1) -
sse42
x86_64's Streaming SIMD Extensions 4.2 (SSE4.2) - tbm Trailing Bit Manipulation (TBM) instruction set.
-
xsave
x86_64'sxsaveandxsaveopttarget feature intrinsics
Structs
- __tile1024i A tile register, used by AMX instructions.
Functions
-
__tile_cmmimfp16ps
Perform matrix multiplication of two tiles containing complex elements and accumulate the results into a packed single precision tile.
Each dword element in input tiles a and b is interpreted as a complex number with FP16 real part and FP16 imaginary part.
Calculates the imaginary part of the result. For each possible combination of (row of a, column of b),
it performs a set of multiplication and accumulations on all corresponding complex numbers (one from a and one from b).
The imaginary part of the a element is multiplied with the real part of the corresponding b element, and the real part of
the a element is multiplied with the imaginary part of the corresponding b elements. The two accumulated results are added,
and then accumulated into the corresponding row and column of dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_cmmrlfp16ps
Perform matrix multiplication of two tiles containing complex elements and accumulate the results into a packed single precision tile.
Each dword element in input tiles a and b is interpreted as a complex number with FP16 real part and FP16 imaginary part.
Calculates the real part of the result. For each possible combination of (row of a, column of b),
it performs a set of multiplication and accumulations on all corresponding complex numbers (one from a and one from b).
The real part of the a element is multiplied with the real part of the corresponding b element, and the negated imaginary part of
the a element is multiplied with the imaginary part of the corresponding b elements.
The two accumulated results are added, and then accumulated into the corresponding row and column of dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_cvtrowd2ps
Moves a row from a tile register to a zmm register, converting the packed 32-bit signed integer
elements to packed single-precision (32-bit) floating-point elements.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_cvtrowps2bf16h
Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit)
floating-point elements to packed BF16 (16-bit) floating-point elements. The resulting
16-bit elements are placed in the high 16-bits within each 32-bit element of the returned vector.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_cvtrowps2bf16l
Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit)
floating-point elements to packed BF16 (16-bit) floating-point elements. The resulting
16-bit elements are placed in the low 16-bits within each 32-bit element of the returned vector.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_cvtrowps2phh
Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit)
floating-point elements to packed half-precision (16-bit) floating-point elements. The resulting
16-bit elements are placed in the high 16-bits within each 32-bit element of the returned vector.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_cvtrowps2phl
Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit)
floating-point elements to packed half-precision (16-bit) floating-point elements. The resulting
16-bit elements are placed in the low 16-bits within each 32-bit element of the returned vector.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbf16ps
Compute dot-product of FP16 (16-bit) floating-point pairs in tiles a and b,
accumulating the intermediate single-precision (32-bit) floating-point elements
with elements in dst, and store the 32-bit result back to tile dst. The shape of the tile
is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbf8ps
Compute dot-product of BF8 (8-bit E5M2) floating-point elements in tile a and BF8 (8-bit E5M2)
floating-point elements in tile b, accumulating the intermediate single-precision
(32-bit) floating-point elements with elements in dst, and store the 32-bit result
back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbhf8ps
Compute dot-product of BF8 (8-bit E5M2) floating-point elements in tile a and HF8
(8-bit E4M3) floating-point elements in tile b, accumulating the intermediate single-precision
(32-bit) floating-point elements with elements in dst, and store the 32-bit result
back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbssd
Compute dot-product of bytes in tiles with a source/destination accumulator.
Multiply groups of 4 adjacent pairs of signed 8-bit integers in a with corresponding
signed 8-bit integers in b, producing 4 intermediate 32-bit results.
Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbsud
Compute dot-product of bytes in tiles with a source/destination accumulator.
Multiply groups of 4 adjacent pairs of signed 8-bit integers in a with corresponding
unsigned 8-bit integers in b, producing 4 intermediate 32-bit results.
Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbusd
Compute dot-product of bytes in tiles with a source/destination accumulator.
Multiply groups of 4 adjacent pairs of unsigned 8-bit integers in a with corresponding
signed 8-bit integers in b, producing 4 intermediate 32-bit results.
Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpbuud
Compute dot-product of bytes in tiles with a source/destination accumulator.
Multiply groups of 4 adjacent pairs of unsigned 8-bit integers in a with corresponding
unsigned 8-bit integers in b, producing 4 intermediate 32-bit results.
Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dpfp16ps
Compute dot-product of FP16 (16-bit) floating-point pairs in tiles a and b,
accumulating the intermediate single-precision (32-bit) floating-point elements
with elements in dst, and store the 32-bit result back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dphbf8ps
Compute dot-product of HF8 (8-bit E4M3) floating-point elements in tile a and BF8
(8-bit E5M2) floating-point elements in tile b, accumulating the intermediate single-precision
(32-bit) floating-point elements with elements in dst, and store the 32-bit result
back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_dphf8ps
Compute dot-product of HF8 (8-bit E4M3) floating-point elements in tile a and HF8 (8-bit E4M3)
floating-point elements in tile b, accumulating the intermediate single-precision
(32-bit) floating-point elements with elements in dst, and store the 32-bit result
back to tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_loadd
Load tile rows from memory specified by base address and stride into destination tile dst. The shape
of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_loaddrs
Load tile rows from memory specified by base address and stride into destination tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. Additionally, this intrinsic indicates the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future. -
__tile_movrow
Moves one row of tile data into a zmm vector register
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_stored
Store the tile specified by src to memory specified by base address and stride. The shape of the tile
is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. -
__tile_stream_loadd
Load tile rows from memory specified by base address and stride into destination tile dst. The shape
of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. This intrinsic provides a hint to the implementation that the data will likely not be reused in the near future and the data caching can be optimized accordingly. -
__tile_stream_loaddrs
Load tile rows from memory specified by base address and stride into destination tile dst.
The shape of the tile is specified in the struct of
__tile1024i. The register of the tile is allocated by the compiler. Provides a hint to the implementation that the data would be reused but does not need to be resident in the nearest cache levels. Additionally, this intrinsic indicates the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future. -
__tile_zero
Zero the tile specified by
dst. The shape of the tile is specified in the struct of__tile1024i. The register of the tile is allocated by the compiler. -
_addcarry_u64
Adds unsigned 64-bit integers
aandbwith unsigned 8-bit carry-inc_in(carry or overflow flag), and store the unsigned 64-bit result inout, and the carry-out is returned (carry or overflow flag). -
_addcarryx_u64
Adds unsigned 64-bit integers
aandbwith unsigned 8-bit carry-inc_in(carry or overflow flag), and store the unsigned 64-bit result inout, and the carry-out is returned (carry or overflow flag). -
_andn_u64
Bitwise logical
ANDof invertedawithb. -
_bextr2_u64
Extracts bits of
aspecified bycontrolinto the least significant bits of the result. -
_bextr_u64
Extracts bits in range [
start,start+length) fromainto the least significant bits of the result. -
_bextri_u64
Extracts bits of
aspecified bycontrolinto the least significant bits of the result. -
_bittest64
Returns the bit in position
bof the memory addressed byp. -
_bittestandcomplement64
Returns the bit in position
bof the memory addressed byp, then inverts that bit. -
_bittestandreset64
Returns the bit in position
bof the memory addressed byp, then resets that bit to0. -
_bittestandset64
Returns the bit in position
bof the memory addressed byp, then sets the bit to1. -
_blcfill_u64
Clears all bits below the least significant zero bit of
x. -
_blci_u64
Sets all bits of
xto 1 except for the least significant zero bit. -
_blcic_u64
Sets the least significant zero bit of
xand clears all other bits. -
_blcmsk_u64
Sets the least significant zero bit of
xand clears all bits above that bit. -
_blcs_u64
Sets the least significant zero bit of
x. -
_blsfill_u64
Sets all bits of
xbelow the least significant one. - _blsi_u64 Extracts lowest set isolated bit.
- _blsic_u64 Clears least significant bit and sets all other bits.
- _blsmsk_u64 Gets mask up to lowest set bit.
-
_blsr_u64
Resets the lowest set bit of
x. - _bswap64 Returns an integer with the reversed byte order of x
-
_bzhi_u64
Zeroes higher bits of
a>=index. - _cvtmask64_u64 Convert 64-bit mask a into an integer value, and store the result in dst.
- _cvtu64_mask64 Convert integer value a into an 64-bit mask, and store the result in k.
-
_fxrstor64
Restores the
XMM,MMX,MXCSR, andx87FPU registers from the 512-byte-long 16-byte-aligned memory regionmem_addr. -
_fxsave64
Saves the
x87FPU,MMXtechnology,XMM, andMXCSRregisters to the 512-byte-long 16-byte-aligned memory regionmem_addr. - _lzcnt_u64 Counts the leading most significant zero bits.
-
_mm256_extract_epi64
Extracts a 64-bit integer from
a, selected withINDEX. -
_mm256_insert_epi64
Copies
ato result, and insert the 64-bit integeriinto result at the location specified byindex. -
_mm_crc32_u64
Starting with the initial value in
crc, return the accumulated CRC32-C value for unsigned 64-bit integerv. - _mm_cvt_roundi64_sd Convert the signed 64-bit integer b to a double-precision (64-bit) floating-point element, store the result in the lower element of dst, and copy the upper element from a to the upper element of dst. Rounding is done according to the rounding[3:0] parameter, which can be one of:\
- _mm_cvt_roundi64_sh Convert the signed 64-bit integer b to a half-precision (16-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst.
- _mm_cvt_roundi64_ss Convert the signed 64-bit integer b to a single-precision (32-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst. Rounding is done according to the rounding[3:0] parameter, which can be one of:\
-
_mm_cvt_roundsd_i64
Convert the lower double-precision (64-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ -
_mm_cvt_roundsd_si64
Convert the lower double-precision (64-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ -
_mm_cvt_roundsd_u64
Convert the lower double-precision (64-bit) floating-point element in a to an unsigned 64-bit integer, and store the result in dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ - _mm_cvt_roundsh_i64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
- _mm_cvt_roundsh_u64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit unsigned integer, and store the result in dst.
- _mm_cvt_roundsi64_sd Convert the signed 64-bit integer b to a double-precision (64-bit) floating-point element, store the result in the lower element of dst, and copy the upper element from a to the upper element of dst. Rounding is done according to the rounding[3:0] parameter, which can be one of:\
- _mm_cvt_roundsi64_ss Convert the signed 64-bit integer b to a single-precision (32-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst. Rounding is done according to the rounding[3:0] parameter, which can be one of:\
-
_mm_cvt_roundss_i64
Convert the lower single-precision (32-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ -
_mm_cvt_roundss_si64
Convert the lower single-precision (32-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ -
_mm_cvt_roundss_u64
Convert the lower single-precision (32-bit) floating-point element in a to an unsigned 64-bit integer, and store the result in dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ -
_mm_cvt_roundu64_sd
Convert the unsigned 64-bit integer b to a double-precision (64-bit) floating-point element, store the result in the lower element of dst, and copy the upper element from a to the upper element of dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ - _mm_cvt_roundu64_sh Convert the unsigned 64-bit integer b to a half-precision (16-bit) floating-point element, store the result in the lower element of dst, and copy the upper 1 packed elements from a to the upper elements of dst.
-
_mm_cvt_roundu64_ss
Convert the unsigned 64-bit integer b to a single-precision (32-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst.
Rounding is done according to the rounding[3:0] parameter, which can be one of:\ - _mm_cvti64_sd Convert the signed 64-bit integer b to a double-precision (64-bit) floating-point element, store the result in the lower element of dst, and copy the upper element from a to the upper element of dst.
- _mm_cvti64_sh Convert the signed 64-bit integer b to a half-precision (16-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst.
- _mm_cvti64_ss Convert the signed 64-bit integer b to a single-precision (32-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst.
- _mm_cvtsd_i64 Convert the lower double-precision (64-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
- _mm_cvtsd_si64 Converts the lower double-precision (64-bit) floating-point element in a to a 64-bit integer.
-
_mm_cvtsd_si64x
Alias for
_mm_cvtsd_si64 - _mm_cvtsd_u64 Convert the lower double-precision (64-bit) floating-point element in a to an unsigned 64-bit integer, and store the result in dst.
- _mm_cvtsh_i64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
- _mm_cvtsh_u64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit unsigned integer, and store the result in dst.
-
_mm_cvtsi128_si64
Returns the lowest element of
a. -
_mm_cvtsi128_si64x
Returns the lowest element of
a. -
_mm_cvtsi64_sd
Returns
awith its lower element replaced bybafter converting it to anf64. -
_mm_cvtsi64_si128
Returns a vector whose lowest element is
aand all higher elements are0. -
_mm_cvtsi64_ss
Converts a 64 bit integer to a 32 bit float. The result vector is the input
vector
awith the lowest 32 bit float replaced by the converted integer. -
_mm_cvtsi64x_sd
Returns
awith its lower element replaced bybafter converting it to anf64. -
_mm_cvtsi64x_si128
Returns a vector whose lowest element is
aand all higher elements are0. - _mm_cvtss_i64 Convert the lower single-precision (32-bit) floating-point element in a to a 64-bit integer, and store the result in dst.
- _mm_cvtss_si64 Converts the lowest 32 bit float in the input vector to a 64 bit integer.
- _mm_cvtss_u64 Convert the lower single-precision (32-bit) floating-point element in a to an unsigned 64-bit integer, and store the result in dst.
-
_mm_cvtt_roundsd_i64
Convert the lower double-precision (64-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
Exceptions can be suppressed by passing _MM_FROUND_NO_EXC in the sae parameter. -
_mm_cvtt_roundsd_si64
Convert the lower double-precision (64-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
Exceptions can be suppressed by passing _MM_FROUND_NO_EXC in the sae parameter. -
_mm_cvtt_roundsd_u64
Convert the lower double-precision (64-bit) floating-point element in a to an unsigned 64-bit integer with truncation, and store the result in dst.
Exceptions can be suppressed by passing _MM_FROUND_NO_EXC in the sae parameter. - _mm_cvtt_roundsh_i64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
- _mm_cvtt_roundsh_u64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit unsigned integer with truncation, and store the result in dst.
-
_mm_cvtt_roundss_i64
Convert the lower single-precision (32-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
Exceptions can be suppressed by passing _MM_FROUND_NO_EXC in the sae parameter. -
_mm_cvtt_roundss_si64
Convert the lower single-precision (32-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
Exceptions can be suppressed by passing _MM_FROUND_NO_EXC in the sae parameter. -
_mm_cvtt_roundss_u64
Convert the lower single-precision (32-bit) floating-point element in a to an unsigned 64-bit integer with truncation, and store the result in dst.
Exceptions can be suppressed by passing _MM_FROUND_NO_EXC in the sae parameter. - _mm_cvttsd_i64 Convert the lower double-precision (64-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
-
_mm_cvttsd_si64
Converts the lower double-precision (64-bit) floating-point element in
ato a 64-bit integer with truncation. -
_mm_cvttsd_si64x
Alias for
_mm_cvttsd_si64 - _mm_cvttsd_u64 Convert the lower double-precision (64-bit) floating-point element in a to an unsigned 64-bit integer with truncation, and store the result in dst.
- _mm_cvttsh_i64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
- _mm_cvttsh_u64 Convert the lower half-precision (16-bit) floating-point element in a to a 64-bit unsigned integer with truncation, and store the result in dst.
- _mm_cvttss_i64 Convert the lower single-precision (32-bit) floating-point element in a to a 64-bit integer with truncation, and store the result in dst.
- _mm_cvttss_si64 Converts the lowest 32 bit float in the input vector to a 64 bit integer with truncation.
- _mm_cvttss_u64 Convert the lower single-precision (32-bit) floating-point element in a to an unsigned 64-bit integer with truncation, and store the result in dst.
- _mm_cvtu64_sd Convert the unsigned 64-bit integer b to a double-precision (64-bit) floating-point element, store the result in the lower element of dst, and copy the upper element from a to the upper element of dst.
- _mm_cvtu64_sh Convert the unsigned 64-bit integer b to a half-precision (16-bit) floating-point element, store the result in the lower element of dst, and copy the upper 1 packed elements from a to the upper elements of dst.
- _mm_cvtu64_ss Convert the unsigned 64-bit integer b to a single-precision (32-bit) floating-point element, store the result in the lower element of dst, and copy the upper 3 packed elements from a to the upper elements of dst.
-
_mm_extract_epi64
Extracts an 64-bit integer from
aselected withIMM1 -
_mm_insert_epi64
Returns a copy of
awith the 64-bit integer fromiinserted at a location specified byIMM1. - _mm_stream_si64 Stores a 64-bit integer value in an 8-byte aligned memory location. To minimize caching, the data is flagged as non-temporal (unlikely to be used again soon).
- _mm_tzcnt_64 Counts the number of trailing least significant zero bits.
- _movrs_i16 Moves a 16-bit word from the source to the destination, with an indication that the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future.
- _movrs_i32 Moves a 32-bit doubleword from the source to the destination, with an indication that the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future.
- _movrs_i64 Moves a 64-bit quadword from the source to the destination, with an indication that the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future.
- _movrs_i8 Moves a byte from the source to the destination, with an indication that the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future.
- _mulx_u64 Unsigned multiply without affecting flags.
-
_pdep_u64
Scatter contiguous low order bits of
ato the result at the positions specified by themask. -
_pext_u64
Gathers the bits of
xspecified by themaskinto the contiguous low order bit positions of the result. - _popcnt64 Counts the bits that are set.
- _rdrand64_step Read a hardware generated 64-bit random value and store the result in val. Returns 1 if a random value was generated, and 0 otherwise.
- _rdseed64_step Read a 64-bit NIST SP800-90B and SP800-90C compliant random value and store in val. Return 1 if a random value was generated, and 0 otherwise.
-
_subborrow_u64
Adds unsigned 64-bit integers
aandbwith unsigned 8-bit carry-inc_in. (carry or overflow flag), and store the unsigned 64-bit result inout, and the carry-out is returned (carry or overflow flag). -
_t1mskc_u64
Clears all bits below the least significant zero of
xand sets all other bits. - _tile_cmmimfp16ps Perform matrix multiplication of two tiles containing complex elements and accumulate the results into a packed single precision tile. Each dword element in input tiles a and b is interpreted as a complex number with FP16 real part and FP16 imaginary part. Calculates the imaginary part of the result. For each possible combination of (row of a, column of b), it performs a set of multiplication and accumulations on all corresponding complex numbers (one from a and one from b). The imaginary part of the a element is multiplied with the real part of the corresponding b element, and the real part of the a element is multiplied with the imaginary part of the corresponding b elements. The two accumulated results are added, and then accumulated into the corresponding row and column of dst.
- _tile_cmmrlfp16ps Perform matrix multiplication of two tiles containing complex elements and accumulate the results into a packed single precision tile. Each dword element in input tiles a and b is interpreted as a complex number with FP16 real part and FP16 imaginary part. Calculates the real part of the result. For each possible combination of (row of a, column of b), it performs a set of multiplication and accumulations on all corresponding complex numbers (one from a and one from b). The real part of the a element is multiplied with the real part of the corresponding b element, and the negated imaginary part of the a element is multiplied with the imaginary part of the corresponding b elements. The two accumulated results are added, and then accumulated into the corresponding row and column of dst.
- _tile_cvtrowd2ps Moves a row from a tile register to a zmm register, converting the packed 32-bit signed integer elements to packed single-precision (32-bit) floating-point elements.
- _tile_cvtrowd2psi Moves a row from a tile register to a zmm register, converting the packed 32-bit signed integer elements to packed single-precision (32-bit) floating-point elements.
- _tile_cvtrowps2bf16h Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed BF16 (16-bit) floating-point elements. The resulting 16-bit elements are placed in the high 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2bf16hi Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed BF16 (16-bit) floating-point elements. The resulting 16-bit elements are placed in the high 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2bf16l Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed BF16 (16-bit) floating-point elements. The resulting 16-bit elements are placed in the low 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2bf16li Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed BF16 (16-bit) floating-point elements. The resulting 16-bit elements are placed in the low 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2phh Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed half-precision (16-bit) floating-point elements. The resulting 16-bit elements are placed in the high 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2phhi Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed half-precision (16-bit) floating-point elements. The resulting 16-bit elements are placed in the high 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2phl Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed half-precision (16-bit) floating-point elements. The resulting 16-bit elements are placed in the low 16-bits within each 32-bit element of the returned vector.
- _tile_cvtrowps2phli Moves a row from a tile register to a zmm register, converting the packed single-precision (32-bit) floating-point elements to packed half-precision (16-bit) floating-point elements. The resulting 16-bit elements are placed in the low 16-bits within each 32-bit element of the returned vector.
- _tile_dpbf16ps Compute dot-product of BF16 (16-bit) floating-point pairs in tiles a and b, accumulating the intermediate single-precision (32-bit) floating-point elements with elements in dst, and store the 32-bit result back to tile dst.
- _tile_dpbf8ps Compute dot-product of BF8 (8-bit E5M2) floating-point elements in tile a and BF8 (8-bit E5M2) floating-point elements in tile b, accumulating the intermediate single-precision (32-bit) floating-point elements with elements in dst, and store the 32-bit result back to tile dst.
- _tile_dpbhf8ps Compute dot-product of BF8 (8-bit E5M2) floating-point elements in tile a and HF8 (8-bit E4M3) floating-point elements in tile b, accumulating the intermediate single-precision (32-bit) floating-point elements with elements in dst, and store the 32-bit result back to tile dst.
- _tile_dpbssd Compute dot-product of bytes in tiles with a source/destination accumulator. Multiply groups of 4 adjacent pairs of signed 8-bit integers in a with corresponding signed 8-bit integers in b, producing 4 intermediate 32-bit results. Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
- _tile_dpbsud Compute dot-product of bytes in tiles with a source/destination accumulator. Multiply groups of 4 adjacent pairs of signed 8-bit integers in a with corresponding unsigned 8-bit integers in b, producing 4 intermediate 32-bit results. Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
- _tile_dpbusd Compute dot-product of bytes in tiles with a source/destination accumulator. Multiply groups of 4 adjacent pairs of unsigned 8-bit integers in a with corresponding signed 8-bit integers in b, producing 4 intermediate 32-bit results. Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
- _tile_dpbuud Compute dot-product of bytes in tiles with a source/destination accumulator. Multiply groups of 4 adjacent pairs of unsigned 8-bit integers in a with corresponding unsigned 8-bit integers in b, producing 4 intermediate 32-bit results. Sum these 4 results with the corresponding 32-bit integer in dst, and store the 32-bit result back to tile dst.
- _tile_dpfp16ps Compute dot-product of FP16 (16-bit) floating-point pairs in tiles a and b, accumulating the intermediate single-precision (32-bit) floating-point elements with elements in dst, and store the 32-bit result back to tile dst.
- _tile_dphbf8ps Compute dot-product of HF8 (8-bit E4M3) floating-point elements in tile a and BF8 (8-bit E5M2) floating-point elements in tile b, accumulating the intermediate single-precision (32-bit) floating-point elements with elements in dst, and store the 32-bit result back to tile dst.
- _tile_dphf8ps Compute dot-product of HF8 (8-bit E4M3) floating-point elements in tile a and HF8 (8-bit E4M3) floating-point elements in tile b, accumulating the intermediate single-precision (32-bit) floating-point elements with elements in dst, and store the 32-bit result back to tile dst.
-
_tile_loadconfig
Load tile configuration from a 64-byte memory location specified by
mem_addr. The tile configuration format is specified below, and includes the tile type pallette, the number of bytes per row, and the number of rows. If the specified pallette_id is zero, that signifies the init state for both the tile config and the tile data, and the tiles are zeroed. Any invalid configurations will result in #GP fault. -
_tile_loadd
Load tile rows from memory specified by base address and stride into destination tile dst using the tile configuration previously configured via
_tile_loadconfig. -
_tile_loaddrs
Load tile rows from memory specified by base address and stride into destination tile dst
using the tile configuration previously configured via
_tile_loadconfig. Additionally, this intrinsic indicates the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future. - _tile_movrow Moves one row of tile data into a zmm vector register
- _tile_movrowi Moves one row of tile data into a zmm vector register
- _tile_release Release the tile configuration to return to the init state, which releases all storage it currently holds.
-
_tile_storeconfig
Stores the current tile configuration to a 64-byte memory location specified by
mem_addr. The tile configuration format is as specified in_tile_loadconfig, and includes the tile type pallette, the number of bytes per row, and the number of rows. If tiles are not configured, all zeroes will be stored to memory. -
_tile_stored
Store the tile specified by src to memory specified by base address and stride using the tile configuration previously configured via
_tile_loadconfig. -
_tile_stream_loadd
Load tile rows from memory specified by base address and stride into destination tile dst using the tile configuration
previously configured via
_tile_loadconfig. This intrinsic provides a hint to the implementation that the data will likely not be reused in the near future and the data caching can be optimized accordingly. -
_tile_stream_loaddrs
Load tile rows from memory specified by base address and stride into destination tile dst
using the tile configuration previously configured via
_tile_loadconfig. Provides a hint to the implementation that the data would be reused but does not need to be resident in the nearest cache levels. Additionally, this intrinsic indicates the source memory location is likely to become read-shared by multiple processors, i.e., read in the future by at least one other processor before it is written, assuming it is ever written in the future. -
_tile_zero
Zero the tile specified by
tdest. - _tzcnt_u64 Counts the number of trailing least significant zero bits.
-
_tzmsk_u64
Sets all bits below the least significant one of
xand clears all other bits. -
_xrstor64
Performs a full or partial restore of the enabled processor states using
the state information stored in memory at
mem_addr. -
_xrstors64
Performs a full or partial restore of the enabled processor states using the
state information stored in memory at
mem_addr. -
_xsave64
Performs a full or partial save of the enabled processor states to memory at
mem_addr. -
_xsavec64
Performs a full or partial save of the enabled processor states to memory
at
mem_addr. -
_xsaveopt64
Performs a full or partial save of the enabled processor states to memory at
mem_addr. -
_xsaves64
Performs a full or partial save of the enabled processor states to memory at
mem_addr - bextri_u64
- cmpxchg16b Compares and exchange 16 bytes (128 bits) of data atomically.
- crc32_64_64
- cvtsd2si64
- cvtss2si64
- cvttsd2si64
- cvttss2si64
- fxrstor64
- fxsave64
- ldtilecfg
- llvm_addcarry_u64
- llvm_subborrow_u64
- movrsdi
- movrshi
- movrsqi
- movrssi
- sttilecfg
- tcmmimfp16ps
- tcmmimfp16ps_internal
- tcmmrlfp16ps
- tcmmrlfp16ps_internal
- tcvtrowd2ps
- tcvtrowd2ps_internal
- tcvtrowd2psi
- tcvtrowps2bf16h
- tcvtrowps2bf16h_internal
- tcvtrowps2bf16hi
- tcvtrowps2bf16l
- tcvtrowps2bf16l_internal
- tcvtrowps2bf16li
- tcvtrowps2phh
- tcvtrowps2phh_internal
- tcvtrowps2phhi
- tcvtrowps2phl
- tcvtrowps2phl_internal
- tcvtrowps2phli
- tdpbf16ps
- tdpbf16ps_internal
- tdpbf8ps
- tdpbf8ps_internal
- tdpbhf8ps
- tdpbhf8ps_internal
- tdpbssd
- tdpbssd_internal
- tdpbsud
- tdpbsud_internal
- tdpbusd
- tdpbusd_internal
- tdpbuud
- tdpbuud_internal
- tdpfp16ps
- tdpfp16ps_internal
- tdphbf8ps
- tdphbf8ps_internal
- tdphf8ps
- tdphf8ps_internal
- tileloadd64
- tileloadd64_internal
- tileloaddrs64
- tileloaddrs64_internal
- tileloaddrst164
- tileloaddrst164_internal
- tileloaddt164
- tileloaddt164_internal
- tilemovrow
- tilemovrow_internal
- tilemovrowi
- tilerelease
- tilestored64
- tilestored64_internal
- tilezero
- tilezero_internal
- vcvtsd2si64
- vcvtsd2usi64
- vcvtsh2si64
- vcvtsh2usi64
- vcvtsi2sd64
- vcvtsi2ss64
- vcvtsi642sh
- vcvtss2si64
- vcvtss2usi64
- vcvttsd2si64
- vcvttsd2usi64
- vcvttsh2si64
- vcvttsh2usi64
- vcvttss2si64
- vcvttss2usi64
- vcvtusi2sd64
- vcvtusi2ss64
- vcvtusi642sh
- x86_bmi2_bzhi_64
- x86_bmi2_pdep_64
- x86_bmi2_pext_64
- x86_bmi_bextr_64
- x86_rdrand64_step
- x86_rdseed64_step
- xrstor64
- xrstors64
- xsave64
- xsavec64
- xsaveopt64
- xsaves64