Test lanes — what runs where, and why¶
Pulp runs its test suite in a few distinct lanes. Knowing which lane a test lands in — and how to route a new test — is the difference between a fast, trustworthy required gate and one that flakes on unrelated work. This is the single source of truth for that model.
The lanes¶
| Lane | Trigger | Gates the PR? | Builds examples? | What it runs |
|---|---|---|---|---|
Required core gate (macos) |
every PR + every merge group | yes (blocking) | Actions: no; Shipyard: yes until promotion | all core tests except the validation, slow, performance, bench, and quality-lab labels; an unchanged exact PR merge tree may reuse its artifact-bound result after protected-base verification |
Example-validation (example-validation) |
PRs touching examples/**, state/format headers, core CMake, or shared dependency infrastructure |
advisory pending promotion (see status below) | yes — Linux + macOS | Linux compiles every example artifact; hosted macOS runs auval + built-in CLAP dlopen checks; pluginval/clap-validator require an operator-dispatched advisory image |
API contracts (api-contracts) |
every PR + every merge group | advisory pending promotion (see below) | no | the Doxygen strict pass over the catalogued public headers, ~3 s of work |
| Nightly full build | schedule (nightly) | no — informational | yes | everything, including all five excluded label groups; results eyeballed, build failures file an issue |
| cross-platform-check | per PR (x86-64 Linux, arm64 Linux, x86-64 Windows) | advisory | no | core tests, excludes validation + slow only — so the timing group does run here, off the reference platform |
The required gate is serialized on self-hosted macOS runners and takes
about 40 min (median 39.8 min of wall-clock over 45 sampled pull_request
runs). Keeping it
lean is why the two label groups below are excluded from it.
Where the required gate's wall-clock goes¶
Measured per-step medians over the same sample, across the 40 runs that allocated a native macOS build (the other 5 resolved to the no-native-build alias and are excluded here, since they run no build or test step). The step medians sum to 39.5 min, slightly under the 39.8 min job median — the difference is per-run scheduling overhead, not a missing step:
| Step | Median | Share |
|---|---|---|
Build |
21.3 min | 54 % |
Test (non-Windows) |
14.8 min | 37 % |
Configure |
1.3 min | 3 % |
| everything else (checkout, brew, ccache install, bootstrap, receipts) | ~2.1 min | 5 % |
Two properties of that profile are worth knowing before optimizing it. Build
already runs against a warm build dir at a 92-94 % ccache hit rate, so it is
dominated by linking rather than by recompilation. Test is not uniformly
parallel: under -j8 the run reaches concurrency 5-8 for its first ~4 min and
then drops to a single test at a time for a median 10.0 min (n=6), because
RUN_SERIAL and PROCESSORS 8 tests can only be scheduled alone and CTest
defers them until the parallel queue drains. pulp-browser-capture-node-integration
alone is ~6 min of that tail and is always the last test to finish; it is
RUN_SERIAL deliberately, because a real-Chrome CDP screenshot crosses its
bounded deadline when unrelated CTest work shares the machine. The serial set,
not the test count, is what sets the floor on the test phase.
For a merge group whose base, head, tree, policy, toolchain record, and tested
artifact identities exactly match a successful PR receipt, the protected-base
verifier derives a new merge-group-bound decision and the native repetition is
omitted. Any missing or changed binding retains the ordinary full gate.
The label taxonomy (how routing works)¶
Routing is driven entirely by CTest LABELS, set in each test's
set_tests_properties(... PROPERTIES LABELS "..."):
validation— a real-host format-validator (pluginval-*,auval-*,clap-dlopen-*). Every user of this label lives underexamples/— it is, in practice, "an example plugin's runtime validation." Slow (apluginvalrun is ~25-30 s) and flaky under concurrent load. Excluded from the required gate; reported by the advisoryexample-validationlane and also run nightly. They do not block merges until that context is promoted.slow— a genuinely long test (e.g.cmake-ios-auv3-configure, a ~25-30 min iOS try-compile). Excluded from the ordinary required corpus; run nightly. A slow test that must gate affected changes needs an explicit affected-surface step in the required job.performance/bench/quality-lab— a relative-timing, CPU-budget, or benchmark measurement (e.g. the sampler heritage suite's "Representative chain stays within the shipping CPU budget", which asserts a ratio against an in-run baseline). Excluded from the required gate since 2026-07-21. These are robust to steady load but not to the load variance produced when the Studio runs its two concurrent build VMs: a sibling VM's bursty compile inflates the ratio past threshold and the verdict tracks runner load rather than the code. They still run on push, on the nightly, and oncross-platform-check. A timing test that must gate belongs in a dedicated cap=1 perf lane, not on the merge path.- no special label — a normal unit/integration test. Runs on the required gate. This is where the vast majority of tests belong.
The required gate excludes all three groups with one CTest filter,
--label-exclude "validation|slow|performance|bench|quality-lab" — the same
filter build.yml's PR ctest uses on pull_request, workflow_dispatch, and
merge_group. The other lanes filter differently and deliberately: a push to
main excludes only validation, and cross-platform-check.yml excludes only
validation|slow. Read the lane you mean; the filters are not uniform. It is set in
.shipyard/config.toml ([validation.default],
test =).
Affected slow proofs¶
agent-capability-installed-sdk installs Pulp and builds an independent
consumer for every exported capability and typed binding. It runs as its own
step on the required macOS job, where it measures a median 129 s (n=17,
range 104-174 s); charging that to every unrelated PR made a meaningful share
of the required test phase one irrelevant proof. It carries
slow;agent-capability-installed-sdk and is restored by build.yml only when
the exact diff touches the capability skill, installed manifest/schemas,
capability history, registry/generator, compile projection, CMake target/export
definitions, or their tests. The restored run executes on the required macOS
job and the parallel Linux matrix leg so platform-specific exports retain their
pre-merge proof; even an otherwise skip-safe selected documentation change
allocates those jobs. Relevant changes therefore still fail before merge, while unrelated PRs
and merge groups do not pay its cost. An unknown diff fails closed and runs it.
Read that selector honestly before treating it as narrow: the pattern list in
tools/scripts/classify_changes.py includes a bare *.cmake / **/*.cmake,
so any CMake file anywhere selects the proof — including a
test/cmake/*_tests.cmake manifest edited only to register an unrelated test.
Because "tests ship with fixes" makes such an edit routine, the proof is
selected by roughly half of the open PR population at any time (20 of 41 in one
census), and *.cmake is what selects it in the large majority of those.
Narrowing that glob would be a fail-closed weakening of a deliberately
conservative classifier, so it wants an explicit owner decision about which
test/cmake manifests genuinely reach the installed export surface — several
of them (SDK-consumer and smoke manifests) do.
The explicit restoration is limited to reduced PR, merge-group, and Shipyard
dispatch corpora; unfiltered main/nightly runs already include the proof and do
not run it a second time.
Why example validators are off the required gate¶
An example plugin's pluginval/auval run has real value — a plugin that fails
validation is broken in a real DAW — but it has no business gating an unrelated
core PR. Historically pluginval-SuperConvolver-VST3 (an example) flaked ~30 %
of the time on the required gate and cost unrelated PRs hours (see
planning/friction/2026-07-15-*). Two things follow:
- Compile is checked on relevant changes.
build.yml's requiredmacosActions job configures examples OFF. Shipyard's separate blocking[validation.default]temporarily keepsPULP_BUILD_EXAMPLES=ONuntil the always-reporting context below is promoted to a required check. Theexample-validationworkflow compiles the full examples tree on Linux and macOS whenever an example, watched state/format header, core CMake surface, or shared dependency infrastructure changes, so a failure is visible on the relevant PR. Only the runtime validators are macOS-specific. This remains advisory until the status below is promoted. - Available hosted validation runs on the PR that changes the example. The
example-validationlane (.github/workflows/examples-validation.yml) runs the registeredvalidation-labeled tests whenever a PR touchesexamples/**. Hosted macOS suppliesauvaland the built-in CLAP dlopen checks;pluginvalandclap-validatorrun only on an operator-dispatched isolated advisory image that installs them. It is deliberately not a nightly-only deferral: a broken example validator is reported on the PR that introduced it. The nightly is only a backstop.
example-validation lane status¶
The lane ships not yet in required_status_checks. It always runs and
reports a stable example-validation status (it internally skips the heavy work
on non-examples/** PRs), so it is required-safe — it can be added to branch
protection without the "Expected — waiting for status" dead-lock GitHub imposes
on a paths:-filtered required check. Promote it to required after one green
real-runner run on an examples/** PR. Until then it is visible-but-advisory.
The API-contract lane¶
A public symbol under a catalogued module root (core/timeline/include,
core/music/include, core/timeline_editor/include, core/timeline_view/include)
must carry a doc comment. tools/build-api-docs.sh --contract-only runs Doxygen's
strict pass and tools/scripts/timeline_api_docs_check.py over the result, and
nothing else — about three seconds after checkout.
It has its own workflow rather than a step inside the docs preview build, and the
split is the point. The same check used to run only inside docs-material.yml,
which is not a required context. On 2026-08-16 it detected an undocumented public
typedef, reported FAILURE before the PR merged, and the PR merged anyway; main's
docs build then failed for eight hours and four unrelated PRs carried a red build
none of them caused. A check that can name a main-breaking defect but not prevent it
converts one bad merge into N misleading reds, which teaches everyone to ignore red.
Two properties of the workflow exist solely so it can be promoted to a required
context, and both fail silently if removed — tools/scripts/test_api_contracts_workflow.py
pins them:
- It reports on
merge_group. A required context that does not fire for a queued group leaves the queue waiting on a result that never arrives. - It has no
pathsfilter. GitHub treats a required context that never reports as permanently pending, so a path-filtered required check blocks every PR outside its filter forever. The check is cheap enough to run unconditionally, so it does. (merge_groupdoes not supportpathsat all.)
The published HTML render stays out of this lane deliberately: it is roughly an
order of magnitude more work and produces a preview artifact, not a verdict. It
continues to run in docs-material.yml, which re-checks the contract on its way to
the render. Putting the render back on this lane would repeat the mistake that put
example validators on the required gate.
Its one external dependency is Doxygen from the runner image's apt mirror, and
that install retries. A mirror hiccup is the most common hosted-Linux flake class,
and a lane meant to gate merges cannot fail on one. The version is deliberately
not pinned: build-api-docs.sh documents that CI's Ubuntu package and a
developer's Homebrew build disagree on some diagnostics, so pinning this lane
alone would make one runner image's version the contract while docs-material.yml
and docs-deploy.yml drifted from it. If image drift ever does flip this check,
pin all three together rather than just this one.
Status: advisory until promoted. Until api-contracts is added to main's
required_status_checks, this lane reports the same defect the old one did and is
equally unable to stop it. Promotion is a branch-protection change:
ghapp api -X PATCH repos/Generous-Corp/pulp/branches/main/protection/required_status_checks \
-f 'contexts[]=Enforce version & skill sync' \
-f 'contexts[]=Build + prove + (owner-gated) deploy' \
-f 'contexts[]=Vellum trusted freeze' \
-f 'contexts[]=Vellum freeze' \
-f 'contexts[]=macos' \
-f 'contexts[]=api-contracts'
When a test's premise cannot hold in a lane¶
Labels route tests that are slow or flaky in a lane. A third case is neither: a test whose premise is false there, so it can only ever report a red that means nothing.
The worked example is the trusted-host launch tests.
loaded_runtime_closure_matches_policy() rejects any image mapped into a launched
child that is neither an Apple platform image nor a pinned inventory file — that
rejection is the security property under test. A sanitizer build injects exactly
such an image into every process it produces: clang++ -fsanitize=undefined links
@rpath/libclang_rt.ubsan_osx_dynamic.dylib, which dyld resolves under the Xcode
toolchain, and Apple's clang has no static sanitizer runtime on macOS to avoid it.
So those tests cannot pass under a sanitizer, and the correct response is neither
to retry them nor to relax the production check — relaxing it would delete the
property the test exists to prove.
Express this in the test source, not with a label:
- Guard the affected
TEST_CASEs onPULP_TEST_WITH_SANITIZER, which a target picks up via$<$<BOOL:${PULP_SANITIZER}>:PULP_TEST_WITH_SANITIZER=1>. - Skip with a stated reason:
SKIP("…"), neverSUCCEED,WARNor a barereturn;. Those three leave the Catch2 case passing, so the lane records nothing and the suite's pass count is identical whether the case ran or its precondition vanished.SKIP()is what ctest surfaces as***Skipped(tools/cmake/PulpCatch.cmakesetsSKIP_RETURN_CODE 4on every discovered case) and what the required gate's non-run summary lists.tools/scripts/check_skip_not_pass.pychecks this mechanically. - Keep the explanation in one place —
test/support/control_runtime_closure_sanitizer.hppholds it for this case, including the measured dylib path.
Source-level guarding is deliberate. A label excludes a whole target: for these
files it would have dropped 15 passing tests to silence 7 impossible ones. And
coverage is preserved where it counts — the guarded cases still run, and still
gate, on every non-sanitizer lane including the required macos gate.
Point the guard the right way, and prove it. An inverted guard is silent in
both directions: the sanitizer lane goes red exactly as before, while every other
lane quietly stops exercising the code. pulp-test-control-runtime-closure-sanitizer-guard
asserts whichever direction is true for the build it is compiled into, so each
lane checks its own.
Adding a test — where will it land?¶
- A core unit/integration test → add it with no special label. It runs on the required gate. Keep it fast (< a few seconds) and non-flaky.
- A new example plugin → its
clap-dlopen/auval/pluginvalvalidators should carryLABELS "validation;<format>"(match the existing examples). That automatically keeps them off the required gate and onto the example-validation lane. GivepluginvalaTIMEOUTcomfortably above its real runtime (e.g.120— SuperConvolver runs ~25-30 s; 30 s was too tight and flaked). - A genuinely long test (minutes) →
LABELS "slow", and make sure something (nightly, or a dedicated lane) actually runs it — do not rely on the informational nightly alone if it must be enforced. - A long proof needed only for a bounded source contract → give it a
descriptive label in addition to
slow, add a fail-closed affected-diff classifier, and restore it explicitly in the required job on that surface.
The trap to avoid¶
Labeling a test slow, validation, performance, bench, or quality-lab
removes it from the required gate. If
nothing else runs it as a gate, you have silently disabled it — the nightly
runs it but does not fail on it. Before moving a test off the required gate,
make sure it is enforced somewhere. During the staged rollout,
example-validation reports example-validator failures but remains advisory;
promotion to a required context is what turns that signal into enforcement.
Use a dedicated gating lane for anything that must block before then. "It runs
nightly" is a backstop, not enforcement.
A label is not the only way to leave the gate¶
Labels are the visible exit. The quieter one is an opt-in CMake flag: a
test registered inside if(PULP_ENABLE_<FEATURE>) does not run on a lane that
never sets the flag — it is not skipped, it is never registered, so it appears
in no ctest output at all and no label names it. PULP_ENABLE_SCENE3D defaults
OFF, and for a long time nothing in .github/workflows/ or
.shipyard/config.toml set it, so the whole Renderer3D and scene3d surface ran
nowhere while the required gate stayed green. .github/workflows/scene3d-advisory.yml
is the lane that now covers it.
So when you add a test behind an opt-in flag, or add a flag that gates existing tests, name the lane that sets it. The one-line check:
Zero hits means zero coverage. Pair it with a flag you know is wired — for
example grep -rc "PULP_ENABLE_GPU" .github/workflows/build.yml returns a
non-zero count — so that an empty result reads as "not wired" rather than "my
grep was wrong".
More traps worth knowing before you write the lane — every one of them returns a clean, confident, empty answer rather than an error:
ctest -Ris case-sensitive.-R 'renderer3d|scene3d'selects 144 of the gated tests and silently drops the capitalized Catch2 case names; the character-class form-R '[Rr]enderer3[Dd]|[Ss]cene3[Dd]'selects 205. A regex that matches less than you meant still exits 0.- A selection that matches nothing exits 0. Pass
--no-tests=error, and assert a floor on the selected count as well — the first catches an empty selection, the second catches one that merely shrank. -Ris a regex, so a literal(in a case name is a group. A Catch2 case whose name containsrun()is selected zero times byctest -N -R 'run() clamps'and once by-R 'run\(\) clamps'. Control for it with a pattern you know matches — a bare-R 'LV2'selecting 18 tests proves the instrument works while the specific pattern selects none.- A comma in a Catch2 case name makes that name unusable as a filter, and
the way it fails is worse than a miss. Catch2 splits a test spec on commas,
so
./binary "A, B"matches nothing, printsNo tests ranand exits 2 — while the same binary exits 0 on a comma-free name.confirm_failure.shreads that exit asINCONCLUSIVE — the test already fails before any edit, so a working test reports as one that does not cover its code and the honest next move, rewriting the test, is exactly wrong. Name new cases without commas;\,escapes one in an existing name.
A test whose premise cannot hold in CI¶
No workflow in this repo checks out submodules
(submodules: false throughout .github/workflows/), so the private
planning/ submodule is unavailable to every hosted lane by construction. A
test that reads a file from it can only be excluded, never fixed, on such a
lane — and the exclusion should say so, because a bare test name in an
--exclude-regex reads as a suppressed failure rather than a structural one.
Check the transitive dependency, not just the ctest arguments. Both
scene3d-native-slice-handoff-contract and its negative twin need the plan
file, but only the first names it in its arguments; the second reaches it
through a verifier that hardcodes the path. Excluding only the obvious one
leaves a permanent red.