Function _mm_maskz_mul_round_sh

#[target_feature(enable = "avx512fp16")]
pub fn _mm_maskz_mul_round_sh(k: __mmask8, a: __m128h, b: __m128h, ROUNDING: i32) -> __m128h

Multiply the lower half-precision (16-bit) floating-point elements in a and b, store the result in the lower element of dst, and copy the upper 7 packed elements from a to the upper elements of dst using zeromask k (the element is zeroed out when mask bit 0 is not set). Rounding is done according to the rounding parameter, which can be one of:

Intel's documentation