Command programs¶
Example files: loom/docs/examples/guide/command-programs/
A kernel specializes one device entry point. A command program specializes a reusable subgraph: its kernel launches, immutable parameters, launch counts, and explicit dependency structure. The result keeps the same source, linking, fact, and target model as an individual kernel instead of introducing a second graph language around it.
In this chapter, you will learn:
- how a command-program signature separates specialization inputs from explicit buffer roots and their materialization roles;
- how named parameter content becomes a typed kernel argument;
- how program declarations and launches compose independently packaged subgraphs;
- how
command.serialandcommand.concurrentstate dependency edges; and - what command preparation preserves, specializes, and emits.
The signature separates host decisions from storage¶
command.program.def has
two argument groups:
command.program.def @project_layer(%layer: index, %element_count: index) launch(
%parameters: buffer, %input: buffer, %output: buffer) {
// Reusable command schedule.
command.return
}
The arguments before launch specialize the program. The arguments after
launch are buffer roots supplied through the program ABI:
| Argument group | Becomes known | What it can control |
|---|---|---|
| Specializations | Facts participate during preparation; residual values feed host launch-count evaluation | Exact source control flow and parameter names, residual launch workloads, provider selection, and derived kernel facts. |
| Buffer roots | Roles are fixed during preparation; storage is bound during materialization or issue | Immutable parameter storage or replaceable invocation storage. |
This distinction is stronger than a conventional kernel argument list.
Specializations are host-side program inputs, not push constants read by every
kernel. Their known facts participate in preparation of the complete command
root. An exact %layer can disappear into parameter placement and selected
branches. A residual %element_count can remain as an input to the generated
host launch-count function without entering the command or device ABI.
Schedule topology and parameter identity must be closed after preparation. Residual specialization values may feed pure launch-count expressions, but they cannot leave an unresolved source branch or a dynamic parameter key in the materialized command program.
The launch-binding group is intentionally buffers only. Preparation classifies
roots used by command.parameter as fixed materialization resources and
ordinary roots as rebindable issue-time resources. Residual scalar
specializations are host inputs to launch-count evaluation; they do not
silently become device push constants. A scalar used by varying device
arithmetic must live in device-visible storage today, while an exact scalar
kernel argument may specialize away. The type checker rejects non-buffer
launch bindings so a caller and a command program cannot silently disagree
about which ABI carries a value.
Workloads and specializations are different
A kernel.launch @entry[%count](...) workload configures that kernel's
physical launch without changing its device ABI. A
command.program.launch @program[%variant](...) specialization contributes
facts to the complete reachable subgraph. Exact facts may change which
kernels exist; residual facts may remain in host launch-count evaluation.
Turn named parameter content into typed views¶
The layer example declares the kernel contract it uses, derives a typed view from one parameter root, and launches the kernel:
Source: loom/docs/examples/guide/command-programs/layer.loom
// The declaration records the exact kernel contract used by this module.
kernel.decl @project_block(%element_count: index) launch(%weights: view<8xf32>, %input: buffer, %output: buffer)
// The layer must become exact for parameter selection. Element count may
// remain a residual host value used to evaluate physical launch counts.
command.program.def @project_layer(%layer: index, %element_count: index) launch(%parameters: buffer, %input: buffer, %output: buffer) where [range(%layer, 0, 127), range(%element_count, 1, 8)] {
%weights = command.parameter %parameters, "blk.{}.projection.weight"[%layer] : view<8xf32>
kernel.launch @project_block[%element_count](%weights, %input, %output) : [index](view<8xf32>, buffer, buffer)
command.return
}
command.parameter is a
declarative association. It performs no string lookup, allocation, transfer,
or synchronization at execution time. During command preparation, each {} in
the pattern receives one canonical decimal index substitution. The complete
key and exact view footprint become a parameter requirement attached to the
source buffer root.
The model composition supplies exact layer constants. A %layer value of 12
therefore produces the key blk.12.projection.weight, while the forwarded
%element_count may remain residual. The view<8xf32> result states the exact
32-byte parameter block available to the kernel. %element_count independently
selects a prefix from one through eight, so the view shape does not create one
kernel version per active element count. Command lowering can place that content
in a fixed resource range while %input and %output remain replaceable
issue-time bindings.
The substitutions must become exact nonnegative indices before the program is prepared. The result type must also have an exact byte footprint. These are materialization contracts rather than runtime error cases: a program whose parameter identity or size remains unknown cannot describe an immutable command artifact.
The kernel receives the typed view directly:
Source: loom/docs/examples/guide/command-programs/kernels.loom
// Multiplies up to eight activation elements by layer-specific weights. The
// parameter view is already typed and placed when the kernel is derived for a
// command program.
kernel.def @project_block(%element_count: index) {
%unit = index.constant 1 : index
kernel.launch.config workgroups(%element_count, %unit, %unit) workgroup_size(%unit, %unit, %unit) : index
} launch(%weights: view<8xf32>, %input: buffer, %output: buffer) {
%element = kernel.workgroup.id<x> : index
%base = index.constant 0 : offset
%input_view = buffer.view %input[%base] : buffer -> view<8xf32>
%output_view = buffer.view %output[%base] : buffer -> view<8xf32>
%weight = view.load %weights[%element] : view<8xf32> -> f32
%input_value = view.load %input_view[%element] : view<8xf32> -> f32
%result = scalar.mulf %input_value, %weight : f32
view.store %result, %output_view[%element] : f32, view<8xf32>
kernel.return
}
The view is more informative than an opaque pointer plus an offset. Its root, byte range, element type, and shape remain available while the dependency kernel is derived. The kernel can therefore use the same ordinary view and vector operations as a standalone kernel; it has no command-program-specific device implementation.
Expand model structure from configuration¶
Model-level configuration can determine command topology without turning
per-invocation shapes into global variants. This program resolves
@model.layer_count, fully unrolls a source loop during command preparation,
and uses each exact induction value to select one layer's parameter block:
Source: loom/docs/examples/guide/command-programs/stack.loom
// Layer count is a model-level specialization: command preparation resolves it
// before expanding the loop into a concrete schedule.
config.decl @model.layer_count : %value: index where [range(%value, 1, 256)]
kernel.decl @project_block(%element_count: index) launch(%weights: view<8xf32>, %input: buffer, %output: buffer)
command.program.def public @layer_stack(%element_count: index) launch(%parameters: buffer, %state: buffer) where [range(%element_count, 1, 8)] {
%layer_count = config.get @model.layer_count : index
%first_layer = index.constant 0 : index
%layer_step = index.constant 1 : index
scf.for %layer = [%first_layer to %layer_count step %layer_step] unroll {
%weights = command.parameter %parameters, "blk.{}.projection.weight"[%layer] : view<8xf32>
kernel.launch @project_block[%element_count](%weights, %state, %state) : [index](view<8xf32>, buffer, buffer)
}
command.return
}
The unroll policy is part of the source contract. Command preparation
requires the layer count to resolve, expands the loop into a finite sequence,
and then sees parameter keys such as blk.0.projection.weight and
blk.1.projection.weight. The %element_count specialization remains a
separate value that controls each kernel workload; it is not promoted to
configuration merely because every layer consumes it.
Compose programs through exact contracts¶
Command programs are symbols. A bodyless
command.program.decl
records the exact specialization and binding contract required by one source
module. command.program.launch
provides both groups explicitly at each call site.
The model module depends on the layer archive and composes it in two different ways:
Source: loom/docs/examples/guide/command-programs/model.loom
// The declaration makes the child program contract explicit without choosing
// where its definition is packaged.
command.program.decl @project_layer(%layer: index, %element_count: index) launch(%parameters: buffer, %input: buffer, %output: buffer) where [range(%layer, 0, 127), range(%element_count, 1, 8)]
// Serial composition orders both launches; the intermediate binding carries
// the first result into the second launch. Exact layer IDs specialize parameter
// names while element count remains a root launch-count input.
command.program.def public @two_layer(%element_count: index) launch(%parameters: buffer, %input: buffer, %hidden: buffer, %output: buffer) where [range(%element_count, 1, 8)] {
%first_layer = index.constant 12 : index
%second_layer = index.constant 13 : index
command.serial {
command.program.launch @project_layer[%first_layer, %element_count](%parameters, %input, %hidden) : [index, index](buffer, buffer, buffer)
command.program.launch @project_layer[%second_layer, %element_count](%parameters, %hidden, %output) : [index, index](buffer, buffer, buffer)
}
command.return
}
// Concurrent composition creates no dependency edge between the branches and
// joins both before the program returns.
command.program.def public @two_branch(%element_count: index) launch(%parameters: buffer, %input: buffer, %left_output: buffer, %right_output: buffer) where [range(%element_count, 1, 8)] {
%left_layer = index.constant 20 : index
%right_layer = index.constant 21 : index
command.concurrent {
command.program.launch @project_layer[%left_layer, %element_count](%parameters, %input, %left_output) : [index, index](buffer, buffer, buffer)
command.program.launch @project_layer[%right_layer, %element_count](%parameters, %input, %right_output) : [index, index](buffer, buffer, buffer)
}
command.return
}
The declaration does not keep a definition live and the linker does not guess
which undeclared symbol a call intended. Linking the model archive against
the layer archive resolves the exact declaration; omitting that dependency
leaves an unresolved program symbol.
During command preparation, reachable program launches are clone-inlined in callee-before-caller order. Each selected root becomes a closed command body, while independently selected roots and shared definitions remain ordinary symbols in the linked source. Recursive command composition is rejected because it cannot produce a finite materialized schedule.
This is the same composition model used by motifs and kernels:
kernel library layer program model program
@project_block <- @project_layer <- @two_layer
^ <- @two_branch
|
typed parameter view
Packaging boundaries stay source boundaries. They do not become routing layers in the application and they do not prevent specialization from seeing through the final linked closure.
State dependencies with structured schedules¶
command.serial orders each
child after the preceding child completes. In @two_layer, the first layer
writes %hidden before the second layer consumes it.
command.concurrent adds no
dependency edge between sibling commands and joins all siblings on exit. In
@two_branch, both projections may execute independently, but the command
program cannot return until both have completed.
These regions nest, so larger schedules can express serial chains, fork/join
sections, and pipelines without flattening them into host callbacks. The two
public roots in model.loom are the individual serial and concurrent building
blocks; nesting those same regions preserves each region's entry and exit
contract.
The schedule structure owns ordering. Buffer identity or aliasing does not silently create dependency edges between concurrent siblings. If two commands have a read/write hazard, the source must place them in a serial relationship. This makes the schedule inspectable and keeps dependency analysis out of opaque device ABIs.
Specialize the subgraph as one program¶
Preparing one or more public command roots performs a bounded sequence of compiler operations:
- Selectively link the dependency closure of the requested roots.
- Flatten reachable command-program composition.
- Specialize source control flow, propagate facts, and retain pure residual launch-count expressions.
- Resolve named parameter requirements and explicit schedule waves.
- Derive independently compilable kernel units for reachable launch sites.
- Share equivalent dependency units across roots and assign dense executable slots.
- Lower each root to portable command Low and derive its launch-count function.
The dependency kernels retain their ordinary target compilation paths. A
command root can therefore lower to the portable cmd.core target while one
kernel dependency lowers to AMDGPU ISA and another target configuration lowers
the same source to SPIR-V. Command composition changes the unit of
specialization; it does not create a command-only kernel language.
Equivalent launch sites may share one derived kernel unit when they call the same kernel and the scalar facts crossing their workloads and device ABI match. Different facts produce distinct units when specialization can profit from them. This is how one concise program can cover many shapes and variants without carrying a hand-maintained dispatch table.
The prepared command plan owns three complementary artifacts:
| Artifact | Responsibility |
|---|---|
| Command root | Portable resource bindings, command schedule, and executable slots. |
| Launch-count function | Host evaluation of the dynamic launch tuples required by that root. |
| Dependency units | Selectively linked and specialized kernels compiled through their normal targets. |
Keep the deployment boundary honest¶
The examples in this chapter format, verify, link, and archive as Loom libraries. The compiler also has the target-neutral command preparation and lowering path described above. The installed tools do not yet expose one complete public command that prepares, materializes, binds, and issues a command root, so this guide does not disguise a private runtime integration as an end-user recipe.
That missing command is an ergonomics gap at the deployment boundary, not a gap in the source representation. Kernel-only applications can continue to compile and launch individual entries. Embedders can use command preparation through the compiler integration while the stable installed workflow is being defined.
Keep failures in their owning contract¶
| Symptom | Contract to inspect |
|---|---|
| Source control flow remains in the prepared schedule | Its condition did not specialize to an exact value; only launch-count expressions may remain dynamic. |
| A launch binding is rejected | The command ABI accepts buffer roots only. |
| A parameter key cannot be prepared | A {} substitution did not become an exact nonnegative index. |
| A parameter view has unknown storage size | Its type does not provide an exact byte footprint. |
| A child program remains unresolved | Its module omitted the exact declaration or archive dependency. |
| Two supposedly parallel commands race | The source placed a hazard inside command.concurrent. |
| Composition preparation reports a cycle | Reachable command programs recursively launch one another. |
Source modules and canonical text explains declaration and archive ownership. Facts and specialization shows how configuration, assumptions, and target facts constrain the values used here. Compile artifacts covers the public per-kernel compilation workflow available today.