scf.for¶
Bounded counted loop over an index or offset domain with optional loop-carried state.
The positive step visits lower, lower + step, and subsequent values strictly below the exclusive upper bound. An empty domain returns the initial carried state; otherwise results are the last iteration's yielded state. The unused induction value after the final iteration need not fit the target's address carrier.
The result list defines the recurring type scheme for loop-carried state. A dependent result type may refer to sibling loop results; the initial operands, body arguments, and yielded values instantiate that scheme with their corresponding SSA identities. This permits a loop to carry a view whose extent or layout changes each iteration while every use still names the extent and layout in its own scope.
The optional pipeline(%depth) and unroll(%factor) policies accept independent SSA values, including template arguments and arithmetic on specialized target properties. Pipelining runs before unrolling the requested loop. Full linear unrolling of explicitly annotated mixed descendants can expose their global loads before the enclosing cut. Compile reports retain applied schedules and final resource costs; loom-compile-report suggest proposes evidence-backed comparisons. The per-instance schedule search shows checked candidates, resource cliffs, and controlled measurements.
Operation contract¶
| Property | Value |
|---|---|
| Semantic phase | source structure |
| Traits | ImplicitTerminator(scf.yield) |
| Interfaces | LoopLike |
Signature¶
| Kind | Name | Type | Cardinality | Description |
|---|---|---|---|---|
| Operand | lower_bound |
address |
required | Inclusive index or physical byte-offset lower bound. |
| Operand | upper_bound |
address |
required | Exclusive upper bound in the lower-bound address domain. |
| Operand | step |
address |
required | Positive step in the lower-bound address domain. |
| Operand | iter_args |
any |
variadic | Initial loop-carried state. Dependent types use the identities of these initial operands. |
| Operand | pipeline_depth |
index |
optional | Optional SSA read-ahead depth consumed by pipeline-scf-for before unrolling the requested loop. A positive exact depth counts original iterations independently of the unroll factor; depth one leaves the serial loop. Ordinary reads and their prerequisites run ahead of ordered consumers, with guarded startup and drain preserving the finite domain. With compile-time exact loop bounds, proven global loads can advance across ordered workgroup loads, stores and workgroup-memory barriers. Global or unknown writes and global barriers are rejected. Memory-pure convergent consumers, such as subgroup reductions, also require exact bounds so all participants retain the same phase split. Ordered or convergent consumers cannot supply read prerequisites. Nested units remain intact unless an explicitly requested full linear unroll exposes mixed descendants. Read-only reductions keep their queued result. The reconstructed loops retain their own unroll policy. |
| Operand | unroll_factor |
index |
optional | Optional SSA unroll factor policy consumed by unroll transforms. The factor bounds body cloning and does not require the loop trip count to be static. |
| Result | results |
any |
variadic | Final loop-carried state and recurring type scheme. Dependent types may refer to sibling results. |
| Attribute | unroll_policy |
enum ScfForUnrollPolicy |
optional | Optional bare unroll policy for required full unroll. |
| Attribute | unroll_schedule |
enum ScfForUnrollSchedule |
optional | Optional schedule used when materializing unrolled loop body copies. |
| Region | body |
region | required | Loop body. Carried entry arguments instantiate the result type scheme with body-local identities. Terminated by scf.yield. (single block, terminator scf.yield.) |
Verification constraints¶
SameType(lower_bound, upper_bound, step)IterArgsMatchResults(iter_args, results)YieldCountMatches(body, results)YieldTypesMatch(body, results)
Examples¶
%result = scf.for %iv = [%c0 to %n step %c1](%acc = %init : f32) -> (f32) {
%next = scalar.addf %acc, %acc : f32
scf.yield %next : f32
}
%result = scf.for %iv = [%c0 to %n step %c1](%sum = %initial : f32) -> (f32) pipeline(%depth) {
%value = view.load %input[%iv] : view<[%n]xf32> -> f32
%next = scalar.addf %sum, %value : f32
scf.yield %next : f32
}