Skip to content

scf.for

← scf dialect

Bounded counted loop over an index or offset domain with optional loop-carried state.

The positive step visits lower, lower + step, and subsequent values strictly below the exclusive upper bound. An empty domain returns the initial carried state; otherwise results are the last iteration's yielded state. The unused induction value after the final iteration need not fit the target's address carrier.

The result list defines the recurring type scheme for loop-carried state. A dependent result type may refer to sibling loop results; the initial operands, body arguments, and yielded values instantiate that scheme with their corresponding SSA identities. This permits a loop to carry a view whose extent or layout changes each iteration while every use still names the extent and layout in its own scope.

The optional pipeline(%depth) and unroll(%factor) policies accept independent SSA values, including template arguments and arithmetic on specialized target properties. Pipelining runs before unrolling the requested loop. Full linear unrolling of explicitly annotated mixed descendants can expose their global loads before the enclosing cut. Compile reports retain applied schedules and final resource costs; loom-compile-report suggest proposes evidence-backed comparisons. The per-instance schedule search shows checked candidates, resource cliffs, and controlled measurements.

Operation contract

Property Value
Semantic phase source structure
Traits ImplicitTerminator(scf.yield)
Interfaces LoopLike

Signature

Kind Name Type Cardinality Description
Operand lower_bound address required Inclusive index or physical byte-offset lower bound.
Operand upper_bound address required Exclusive upper bound in the lower-bound address domain.
Operand step address required Positive step in the lower-bound address domain.
Operand iter_args any variadic Initial loop-carried state. Dependent types use the identities of these initial operands.
Operand pipeline_depth index optional Optional SSA read-ahead depth consumed by pipeline-scf-for before unrolling the requested loop. A positive exact depth counts original iterations independently of the unroll factor; depth one leaves the serial loop. Ordinary reads and their prerequisites run ahead of ordered consumers, with guarded startup and drain preserving the finite domain. With compile-time exact loop bounds, proven global loads can advance across ordered workgroup loads, stores and workgroup-memory barriers. Global or unknown writes and global barriers are rejected. Memory-pure convergent consumers, such as subgroup reductions, also require exact bounds so all participants retain the same phase split. Ordered or convergent consumers cannot supply read prerequisites. Nested units remain intact unless an explicitly requested full linear unroll exposes mixed descendants. Read-only reductions keep their queued result. The reconstructed loops retain their own unroll policy.
Operand unroll_factor index optional Optional SSA unroll factor policy consumed by unroll transforms. The factor bounds body cloning and does not require the loop trip count to be static.
Result results any variadic Final loop-carried state and recurring type scheme. Dependent types may refer to sibling results.
Attribute unroll_policy enum ScfForUnrollPolicy optional Optional bare unroll policy for required full unroll.
Attribute unroll_schedule enum ScfForUnrollSchedule optional Optional schedule used when materializing unrolled loop body copies.
Region body region required Loop body. Carried entry arguments instantiate the result type scheme with body-local identities. Terminated by scf.yield. (single block, terminator scf.yield.)

Verification constraints

  • SameType(lower_bound, upper_bound, step)
  • IterArgsMatchResults(iter_args, results)
  • YieldCountMatches(body, results)
  • YieldTypesMatch(body, results)

Examples

scf.for %iv = [%c0 to %n step %c1] {
  scf.yield
}
scf.for %byte_offset = [%zero to %byte_length step %one] {
  scf.yield
}
%result = scf.for %iv = [%c0 to %n step %c1](%acc = %init : f32) -> (f32) {
  %next = scalar.addf %acc, %acc : f32
  scf.yield %next : f32
}
scf.for %iv = [%c0 to %n step %c1] unroll(%factor) {
  scf.yield
}
scf.for %iv = [%c0 to %n step %c1] unroll(%factor) schedule(interleaved) {
  scf.yield
}
%result = scf.for %iv = [%c0 to %n step %c1](%sum = %initial : f32) -> (f32) pipeline(%depth) {
  %value = view.load %input[%iv] : view<[%n]xf32> -> f32
  %next = scalar.addf %sum, %value : f32
  scf.yield %next : f32
}
%result = scf.for %iv = [%c0 to %n step %c1](%sum = %initial : f32) -> (f32) pipeline(%depth) unroll(%factor) schedule(recurrence) {
  %value = view.load %input[%iv] : view<[%n]xf32> -> f32
  %next = scalar.addf %sum, %value : f32
  scf.yield %next : f32
}
%result = scf.for %iv = [%c0 to %n step %c1](%acc = %init : f32) -> (f32) unroll(%factor) schedule(recurrence) {
  %next = scalar.addf %acc, %acc : f32
  scf.yield %next : f32
}