Patch streaming#
When a preprocessing chain allows it, KonfAI reads each patch’s source region straight from disk and never materializes the volume. This page explains what decides whether that happens, and what you control.
Streaming is derived, not configured. Nothing in YAML asks for it. KonfAI reads your preprocessing chain, works out whether each patch’s answer can be computed from a bounded region of the file, and streams when it can. When it cannot, it loads the volume. Both paths give the same patches — exactly for every kind but a resample, which gathers the same samples through a different coordinate frame and lands within the tolerance stated below. They differ in memory and speed.
A 16 GiB uncompressed .mha at patch 64³, batch 2, two workers, under an 8 GiB
memory_budget (below the dataset size, so the run streams) trains at a peak
anonymous RSS of 0.46 GiB, stable across epochs, with VRAM equal to one batch.
The three regimes#
Which one a case takes follows from the workflow’s regime and from the chain itself.
Regime |
When |
Memory held |
|---|---|---|
Cache |
the training default |
every case, resident for the whole run |
Stream |
the predict/eval default, or a budget the dataset exceeds; chain streamable |
one patch |
Buffer |
same triggers, chain not streamable |
a FIFO of |
The cache wins over streaming. A cached case is cut from the resident volume even when its chain would stream; the stream/buffer paths only run on the streaming regime.
Stream and buffer coexist inside one run. The decision is made per case and per augmented copy, not once per dataset, so a chain that streams for one augmentation draw may load the volume for the next.
The regime is not a config key. Training defaults to the cache (epochs re-read
every case); prediction and evaluation default to streaming (one pass). A
memory_budget replaces the default with a measured decision — see below.
What you control#
Two keys under Dataset::
Key |
Default |
Where |
Effect |
|---|---|---|---|
|
|
all three workflows |
Derives the regime from the dataset’s size; an absent key means |
|
|
|
Bounds how many cases stay resident on the buffer path. |
Both are listed in Training configuration.
memory_budget compares the estimated per-rank dataset size against the budget
and picks the regime accordingly, printing its decision once. A bare number is
GiB (24 means 24 GiB), a string may name its unit ("24GB", "32 GiB",
"512mb"), and "auto" — also what an absent key means — offers 80% of the
detected node memory, cgroup limit included, divided by the ranks sharing it.
The budget’s cache side applies to training only: a training dataset that fits caches, one that does not streams — declaring a budget smaller than the dataset is how you force the streaming path. Prediction and evaluation are one-pass workflows: each case is read once, a cache is never re-read, so they always stream whatever the budget says; in evaluation the same budget instead sizes the disjoint patches a too-large case is cut into. Set the budget on the workflow you mean to bound:
Dataset:
memory_budget: auto
Warning
memory_budget is an estimate, not a guarantee. The figure is computed from
file headers alone — prod(shape) × 4 bytes per source group — so it ignores
the dtype you actually stored, any size-changing transform, the augmented copies
of each case, and the transient peak while a case is being built. Treat it as a
coarse switch, not as a memory bound.
Warning
Do not combine shuffle_window with multi-GPU training. Each rank sizes its
sampler from its own shard, so ranks can disagree on the number of batches and
the run hangs in NCCL.
What decides whether a chain streams#
Every transform declares how its output at one voxel depends on its input. That declaration — its patch locality — is what the dispatcher reads to work out which region of the file a patch needs.
Kind |
Meaning |
What KonfAI reads |
|---|---|---|
|
the voxel depends only on itself |
the exact patch |
|
a bounded neighbourhood of radius |
the patch enlarged by |
|
flip or permute |
the index-remapped region |
|
the source region is the target translated |
the region — reading it is the answer |
|
needs whole-volume |
the statistic once from disk, then the exact patch |
|
resample |
the region through the scale mapping, plus an interpolation halo |
|
genuinely needs everything |
the volume — the fallback |
WHOLE_VOLUME is the default. A transform that declares nothing loads the
volume and is correct.
A chain streams when every stage is pointwise or a region kind — HALO,
ORIENTATION, CROP, RESCALE — where GLOBAL_STAT counts as pointwise.
Region stages compose in any number: each stage’s region pulls through the
one before it, down to one bounded read of the stored volume. The chain planned
is the group’s transforms followed by the copy’s own augmentation draw — one
list, so a region transform and a region augmentation compose exactly like two
region transforms.
Five conditions reject streaming:
any
WHOLE_VOLUMEdeclaration;a halo wider than half the read extent on any axis;
a
GLOBAL_STATpreceded by a stage that does not preserve statistics;a
GLOBAL_STATwhose statistic cannot be read from disk;a
RESCALEon a case with noSpacing.
Why [Clip(-200, 400), Standardize()] does not stream#
This is rule 3.
Standardize is GLOBAL_STAT: it needs the volume’s Mean and Std, which
KonfAI reads once from disk. The statistic on disk is the stored volume’s,
while Standardize’s input here is the clipped volume. Clip maps values, so
the two differ. Streaming this chain would standardize every patch by the
pre-Clip statistic, so KonfAI loads the volume instead.
Swap the first stage and the chain streams:
transforms:
Canonical: {}
Normalize: {}
Canonical is ORIENTATION, the one kind that preserves statistics: a flip or a
permute moves voxels without changing any of them, so the multiset of values —
and every statistic over it — is untouched. The stored volume’s Min and Max
are Normalize’s own input, and the chain streams.
The rule generalizes: a GLOBAL_STAT stage streams only when every stage
before it preserves statistics. TensorCast is the one built-in that declares
this for itself, and only for a target that holds every value a volume is read
as: [TensorCast('float32'), Standardize()] streams, while
[TensorCast('uint8'), Standardize()] and [TensorCast('float16'), Standardize()]
do not — a half cast moves 1500.3 onto 1500.0.
How several region stages stream together#
A region stage is a rewrite of which source voxels a target patch needs, and
these rewrites compose. The planner folds them back to front: the target
patch maps through the last stage’s rewrite, that region maps through the stage
before it, and so on down to one region on disk. KonfAI reads that one region
and runs the chain forward over it. [Dilate(1), Gradient()] is two halos that
add, [Canonical(), Permute('2|1|0')] is two reorientations that pull through
each other — both stream from a single bounded read, as does the whole
[Canonical, ResampleToResolution, Padding] reorient-resample-pad pipeline. The
one thing a region stage cannot do is read a statistic of its own input, because
that input never exists whole on disk (rule 3).
Rule 2 is about cost, not correctness. Every patch pays its halo on every side,
so streaming reads prod(1 + 2·halo/extent) times the case’s bytes, where the
extent is the patch or the volume, whichever is smaller. At half the extent that
is 8× in 3D, against the single load streaming avoids. At patch 8, Dilate(4)
streams and Dilate(5) does not.
What each built-in declares#
Kind |
Transforms |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
Augmentations declare per (case, draw) — two copies of the same case can
answer differently. Permute, Flip (when vector_field: false), and Rotate
(when the draw is a quarter turn) are ORIENTATION; ColorTransform and its
subclasses are POINTWISE; Translate is HALO. Scale, Noise, CutOUT,
Mask, and Elastix load the volume.
These transforms load the volume because their answer needs it:
Mask, andClip/Standardizewith amask, read a second full volume that a patch cannot locate itself in.Clipwith percentile bounds needs the whole histogram, as doesHistogramMatching.Savewrites the whole preprocessed volume — but rarely loads it. ASavewhose cache exists becomes the streaming source, and only the transforms after it are planned. ASavewhose cache is missing is materialized slab by slab when the transforms before it stream: each slab of the Save’s own space is read through the composed region plan and region-written, the entry appears only once complete, and the case then streams from it as if the cache had always existed. This runs the prefix once — instead of once per patch per epoch — and lets a global statistic seed after a value-changing stage ([Clip, Save, Standardize]streams;[Clip, Standardize]cannot). Only aSavefed by an unstreamable prefix, or writing to a format without region writes (anything but uncompressedmha,h5, OME-Zarr), still loads the volume.Argmax,Softmax, andSumover a spatialdimreduce across the whole extent.Canonicalon an oblique direction resamples.
Reading regions from disk#
Streaming is only as cheap as the format underneath it.
Backend |
Serves a disk region |
|---|---|
HDF5 |
yes, natively |
OME-Zarr |
yes, chunked ( |
DICOM |
yes, per slice |
SimpleITK |
only uncompressed MetaImage and non-gzipped NIfTI |
A format that cannot serve a region — NRRD, or any compressed file — still
returns the right voxels: it decodes the whole volume for every patch. That
costs speed, never correctness, and KonfAI warns once per format with the
remedy. Convert such datasets to OME-Zarr, HDF5, or uncompressed .mha/.nii.
The same table governs the GLOBAL_STAT seed. On a backend that serves regions,
the statistic is a chunked running pass in float64 — slab by slab, never the
whole volume in RAM. On a format that cannot serve one, reading the statistic
decodes the volume in one go, once per case.
transforms vs patch_transforms#
transforms runs once on the case; patch_transforms runs on each patch after
it is cut. Only POINTWISE and GLOBAL_STAT are admissible per patch, and
KonfAI rejects anything else at config time with the reason and the remedy —
move it to transforms.
A per-patch GLOBAL_STAT means derive the statistic from this patch. To
standardize patches by the volume’s statistic instead, pair a case-level
Standardize(lazy=True) with a per-patch Standardize().
Custom transforms#
Subclass konfai.data.transform.Transform and you inherit the safe default:
your transform reports WHOLE_VOLUME, KonfAI loads the volume, and your patches
are correct without you knowing streaming exists.
Subclassing is required. A class that merely looks like a transform fails
inside the planner with a bare AttributeError rather than a KonfAI error with a
remedy.
To opt in, override patch_locality under three rules:
Read-only. Never write to
cache_attribute. The declaration is made once for the whole case; a write is handed a private copy and lost.No I/O. Read the attribute in hand and nothing else. Whether the declaration can be honoured is the dispatcher’s call.
Total. Answer for any case, including one with no metadata. A missing key returns
WHOLE_VOLUME; it never raises.
ORIENTATION and CROP must also implement stream_region_source, mapping a
target patch to the source region. Declaring a region kind without it raises a
TransformError on the first patch read. HALO and RESCALE need no remap; the
dispatcher derives their regions.
Equivalence#
A streamed patch is byte-identical to the same patch cut from the loaded volume
— border padding included — for POINTWISE, HALO, ORIENTATION, and CROP.
Two cases carry a bounded numerical difference instead.
A GLOBAL_STAT stage seeded from Mean/Std — Standardize — reads its
statistic from a numpy pass over the file while the whole-volume path recomputes
it in torch. Same values, different summation order, so a standardized voxel may
land a few float32 ulp away. A stage seeded from Min/Max — Normalize,
Clip('min', 'max') — is byte-identical: a min and a max have no summation
order to disagree on.
A streamed RESCALE computes interpolation weights in the sub-region’s
coordinate frame, so the deviation scales with the local gradient rather than
with the voxel’s own value: within a few ulp of the volume’s peak on float
volumes, within 1 LSB on integer ones. A nearest-neighbour resample — what a
uint8 label volume gets — uses no weights and stays exact. The Translate
augmentation interpolates on the same terms as RESCALE.
Next steps#
Datasets and groups — the grouped layout,
groups_src, and where patching sitsTraining configuration —
memory_budgetandshuffle_windowin a full configStorage backends & formats — format tokens and the per-backend APIs