vector.dot2f
← vector dialect
Group adjacent two-lane f16 or bf16 products along the last axis and accumulate each pair into an f32 lane using native grouped-dot arithmetic. Intermediate precision, rounding, and subnormal handling follow the selected target's arithmetic contract; an f32 result does not promise the bits of two ordered f32 fused multiply-adds. Target-independent facts, constant folding, and scalar expansion use the reference evaluation: extend each input to f32 and apply scalar.fmaf to the first pair of inputs and accumulator, then to the second pair and partial result. Reference evaluation remains permitted even when native execution differs. Use explicit scalar.fmaf operations or vector.dotf on widened f32 inputs when ordered f32 fused accumulation is required.
Operation contract
| Property |
Value |
| Semantic phase |
executable |
| Traits |
Pure |
Signature
| Kind |
Name |
Type |
Cardinality |
Description |
| Operand |
lhs |
vector |
required |
f16 or bf16 source lanes grouped in pairs along the last axis. |
| Operand |
rhs |
vector |
required |
f16 or bf16 source lanes grouped in pairs along the last axis. |
| Operand |
acc |
vector |
required |
f32 accumulator lanes updated by each two-lane dot product. |
| Result |
result |
vector |
required |
Updated f32 accumulator lanes. |
Verification constraints
HasF16OrBf16Element(lhs)
HasF16OrBf16Element(rhs)
HasF32Element(acc)
SameShape(lhs, rhs)
SameElementType(lhs, rhs)
SameType(acc, result)
LastAxisGroupedBy(lhs, result)
Examples
%r = vector.dot2f %lhs, %rhs, %acc : vector<16xf16>, vector<16xf16>, vector<8xf32>
%r = vector.dot2f %lhs, %rhs, %acc : vector<2x16xbf16>, vector<2x16xbf16>, vector<2x8xf32>