Search by

droost / workflow

chadmandoo

The phased, gated pipeline an agent runs to build or change a Drupal site: plan, code, test, document, complete. Framework-free — the Drupal surface ships with drupal/droost.

0.13.12 2026-10-07 18:46 UTC

README

ci Packagist PHP License

The phased, gated pipeline an agent runs to build or change a Drupal site:

plan → code → test → complete

Each phase has an entry and an exit gate. Pass, and the run advances; fail, and it enters a bounded feedback loop or stops. What "pass" means is configured per repo, in one version-controlled file.

Droost Workflow is the methodology layer. Where droost is what an agent knows about Drupal, this is how it works: the same pipeline, the same levers, whether it runs against a live site or from a plain checkout with no site at all.

Clean-room GPL. A sibling of, not a fork of, any proprietary pipeline.

Status

Built ticket by ticket, and past its skeleton:

Step What State
P6.1 Config spine — the lever file, presets, run state shipped
P6.2 The five phases as a .claude/ pack shipped
P6.3 Gate runner + honest degradation shipped
P6.4 Agentic / interactive modes and the mid-run swap shipped
P6.5 The drush live-site surface and the standalone CLI shipped
P6.6 The MCP surface (optional submodule) shipped
— Hardening: the phase→gate map, engine-counted retries, a coverage gate that can pass shipped
— droost/workflow published on Packagist this release

The lever file

A single repo-root file, droost.workflow.yml, is the source of truth:

mode: agentic                   # agentic | interactive
preset: custom                  # custom | low | medium | high | xhigh | max
enforcement: soft               # hard | soft | off — the hooks, mid-run only
require_run: hard               # hard | soft | off — custom-code edits with NO active run
gates:
  phpcs:          { on: true,  standard: "Drupal,DrupalPractice" }
  # ^ Drupal's standard needs drupal/coder. Where the project is not Drupal —
  # no coder, no docroot, no module/theme/profile at the root — `init` writes
  # PSR12 here, and a level's Drupal default is substituted the same way,
  # with a notice under `deprecations`. Set it yourself and it is yours.
  phpstan:        { on: true,  level: 6 }      # 0-9 | max
  # both static gates accept paths: "web/modules/custom,web/themes/custom" —
  # repo-relative analysis targets. Without paths and without a
  # phpcs.xml/phpstan.neon, the gate is pointed at the project's OWN code:
  # its top-level source directories and root-level files, or a Drupal
  # site's modules/custom and themes/custom — never vendor/, core or contrib.
  # And root: "site" when the Composer project is not the repository root:
  # the tool comes from site/vendor/bin, runs in site/, and discovers
  # site's phpcs.xml(.dist) or phpstan.neon(.dist); findings are still named
  # from the repository. phpunit, coverage and mutation take it too.
  phpunit:        { on: true }
  mutation:       { on: false, msi_min: 0 }
  playwright:     { on: true, required: true }   # a committed spec, at every level
  parity:         { on: true }                 # the site against its source's captured pages
  coverage:       { on: false, min: 0 }
  rendered_check: { on: true }                 # artifacts are truth
  config_clean:   { on: true }                 # a fresh cex produces zero diff
  wiki_fresh:     { on: true }                 # the project's own docs still match the code
  # every gate also takes mode: block | report — report records a failure and
  # advances (never on the mandatory trio)
  # custom:                                    # your own commands as gates
  #   semgrep: { on: true, phase: code, cmd: "semgrep scan --error --quiet" }
  #   suite:   { on: true, phase: test, cmd: "bin/gate.sh FULL", timeout: 900 }
  # contributed:                               # gates enabled MODULES declare (module:<id>)
  #   snyk: { mode: block }                    # on and mode only; the rest is the module's
# baseline: { on: false }                      # strict mode: ignore a committed droost/baseline/
max_gate_retries: 2

It is a file, not Drupal configuration, for four reasons: the agent must be able to read its own levers while the site is mid-build or broken; a plain Claude Code or Codex user reads the same file with no site at all; it is dev tooling and belongs with the code it gates; and it belongs in review, where loosening a gate shows up as a diff. (What each host does and does not get from init is set out under Install — the enforcement hook is Claude Code's alone, and that is a limitation, not a detail.)

Presets — one dial for how hard the workflow verifies

A preset is a base, not an alternative to per-gate control — explicit gates: entries are applied over it, so "max but without Playwright" is one line rather than a fork. Since 2.0 the presets form one graded dial. Two things never move with it: a spec is always written, and the brain (search, the code graph, the wiki as knowledge) is always used — the dial scales what is verified and what artefacts are written, never what the agent must know. Consent (the write wall, require_run) is not in a preset at all.

Level Verification
low basic static checks (phpcs, phpstan 1); no tests; the browser check; no wiki. Seeker off, enforcement soft. The one level whose base turns a mandatory gate off — allowed here and only here, because preset: low is one loud reviewable line, never the default; the gate is recorded off, never passed.
medium the mandatory trio with phpstan 2, the rendered check, a shorter EARS spec, docs in chat. From here up phpunit carries in_diff: a run that changes a class under src/ must change a phpunit test too, or the test phase blocks (tests_in_diff); a test droost's scaffold wrote and nobody changed does not count. And wiki_fresh carries cover_diff: a custom module or theme the run changed must have a wiki page, or the gate fails. (Formerly light, plus both.)
high solid static analysis + unit tests (phpstan 6), no slow tiers; enforcement hard. (The shipped custom gate set, named on the dial, plus in_diff and cover_diff.)
xhigh + coverage and mutation at 60, the front-end trio, phpstan 8.
max everything on, strict — phpstan max, mutation 80, coverage 80 — and tests required to exist: phpunit and playwright carry required: true, so a missing or empty suite is a failure, not a labelled pass. (Formerly factory, plus required.)
custom not a point on the dial: "no opinion — my gates: block is the truth", the spelled-out baseline init writes so choosing it is visible in a diff. Same gates as high, without in_diff or cover_diff.

factory and light still load as aliases of max and medium — the run records the canonical name and a notice says so. (fast was retired in 0.3 and is refused with a pointer.) Switching levels is one line; a run is frozen under the level it started with, so a mid-run edit reshapes the next run.

Off says why. A gate a level turns off is reported off — by preset <level>; a gate the file turns off against its level reads off — by the lever file (preset <level> turns it on) — the same status word, so a reader can tell the dial doing its job from a loosening without opening the file. The reason rides as the result's skip_reason, and every run envelope carries preset, the level the run was held to: "no phpunit result" is readable as a failure or as a low run that never asked for one.

Moving the dial is one command — drush droost:workflow:effort <level> rewrites the preset: line, proves the file still loads (rolling back if not), names the canonical level an alias resolved to, and lists the gate switches the file still spells out (each overrides the dial). It is the operator's command like gate-waive and bypass: it refuses without an interactive terminal and the pack's guard refuses it from the agent's shell — the agent proposes the level, a human sets it. A bare effort only reports, and effort <level> --preview prints the bill — every gate, threshold, the seeker, enforcement and the retry bound the move would change for the next run — without writing, so an operator sees what raising the level applies to existing code before the first run pays for it (round 30 raised a room to max and the first run met 26 legacy phpstan errors), and an agent can ground the level it proposes. A real move prints the same list after it lands; an empty list on a move says why (every touched gate is spelled out in the file). drush droost:workflow:install --preset=<level> writes a fresh file at a level, carrying tuning only (the measured paths) so the level actually drives.

The trio follows the pair. From xhigh up the level turns eslint, stylelint and prettier on, and each is scoped by its own paths. Given none, a trio gate takes phpcs's paths — the project's own code is the same place for both — so a hand-written preset: max with only the PHP pair scoped never lints the repository root. A trio gate given its own paths keeps them. Each also takes config: — the project's own lint config (the file its package.json lint script names), which the gate pins while turning discovery off. Set it on a Drupal docroot: left to discover, eslint's cascade reaches core's scaffolded .eslintrc.json, whose plugins only core's own yarn install provides, and the tool crashes before it reads a file. A crash is reported as ERROR — tool could not run, with the tool's own line and the lever to set: it blocks like a missing tool, and it is never counted as findings — baseline --measure calls such a gate not measurable rather than "0".

A waiver can reopen a terminally failed phase. When a gate spends the retry budget, run refuses and the record says why; recovery used to be reset alone. Now the operator's drush droost:workflow:gate-waive <gate> "<reason>" — signed from a terminal, refused from the agent's shell — also reopens the phase when it covers every gate that killed it: the phase runs again with those gates recorded as waived, never passed, and the reason rides in the report. A waiver on some other gate changes nothing, and the mandatory trio still cannot be waived at all.

The phases are not levers. Since 0.3 every run walks the canonical order, minor changes included — plan → code → test → complete since 0.4 folded document into complete — and the phases: key is deprecated and ignored (with a notice). What varies between heavy and light is the weight each phase carries, never the path.

The mandatory trio is not a lever either. Since 0.4, phpcs, phpstan and phpunit cannot be turned off from the gates: block (the one exception is the low preset's base, which drops phpunit — one loud reviewable line, recorded off, never passed) — they are the toolchain Drupal core itself develops with (exactly what drupal/core-dev ships). Their tuning levers (standard, level, paths) still apply; an on: false, or phpstan's level: off, is recorded as a deprecation notice and superseded. A repo that cannot run one of them yet gets an honest answer instead of a pass: tool missing, config missing (phpunit with no phpunit.xml refuses before spawning anything), or a labeled "nothing to analyse / no tests yet" that hardens itself the moment real code or a first test exists.

The seeker checkpoint holds a green code phase — and completion — until an adversarial inspection is recorded clean. Gates verify rules; the seeker (the pack's workflow-seeker agent) verifies judgment: dead new code, drift from the spec's EARS criteria, coupling the change breaks, weak tests, security smells in the changed code, attempts to defeat the workflow's own discipline. Its scope contract is the diff plus one hop to the changed symbols' consumers — never an audit of the neighbourhood — and its verdict is an exact ledger the engine PARSES (seeker-report, stdin): open CRITICAL or MEDIUM rows hold the run, resolved and carried-with-reason rows release it, and a section with neither rows nor the (no findings) sentinel is an incomplete inspection and refuses. seekers: { on: false } is the one lever, on by default; the hold spends no retry budget.

Enforcement is its own lever, orthogonal to the preset: hard blocks out-of-phase actions while a run is active (editing project files during plan, ending the turn mid-phase), soft warns once per phase, off stands the hooks down. Outside an active run the phase hooks have no opinion — but require_run still stands: a custom-code edit (modules/custom, themes/custom) with no ACTIVE run is blocked (hard, the default even when the key is absent), nudged once (soft), or allowed (off). A finished or failed run counts as no active run — the record is history, not a licence — and the way past the wall is to start a run, or an operator-granted bypass (drush droost:workflow:bypass "<why>", cleared with --off), never the agent's own hand. You may pair factory gates with enforcement: off; not advised, but the lever file is a reviewable diff, and a visible loosening is the honest way to allow it.

Playwright is the npm tier, and it is on at EVERY preset. The gate runs node_modules/.bin/playwright test — committed regression specs, exit code as verdict — with required: true, so an empty suite is a failure rather than a labelled pass.

It is the only gate whose missing binary REPORTS instead of blocking. Every other absent tool is error-tool-missing, which stops the run, and that is right: you asked for the gate and the environment cannot run it. The browser suite is the exception because it is due everywhere, and blocking on an uninstalled binary would wedge every project that has not run npm i -D @playwright/test — an instrument that cannot see refusing to let the run past. The row says what is missing and how to install it, and the moment it IS installed required: true makes this a real wall.

Parity is the site held to the source it rebuilds. A reference is a page of that source (a design, or the site being replaced) read once with vendor/bin/droost-parity capture --source <url> --routes /,/about and kept in the project under droost/parity/: every visible element that carries text, with its box and the computed style that decides how it looks. The gate runs droost-parity judge, reads the same routes on the site with the project's own Playwright, and fails a route whose text is missing or out of order, whose type, colour or place differs past a stated tolerance (every heading's type must match), or whose header and footer differ (D1 to D7, each named in the finding). A page it cannot read is INVALID and fails closed, never a pass. With no reference captured it is a labelled pass that measured nothing, so it is on at every preset; required: true makes a missing reference a failure, and scope: frame judges only the header and footer, for a ticket that builds the frame before the pages. No Playwright, or no Node, REPORTS, as the browser suite does.

A reference may be captured at several widths (--width 1280,390), and every route is judged at each: a narrow width is how a collapsed menu is held to the source's. pages: /,/camps judges those routes whole (page scope) at the first width while every other view keeps scope, so a ticket that rebuilds one page holds it and the frame everywhere at once. The reference is frozen with the run: its digest is taken when the run begins, and the gate fails a reference that has changed since, whatever changed it. Capturing is the operator's, before a run.

Why a spec and not an MCP call. The browser check used to be "the agent called a Playwright MCP tool", counted from the guard's ledger. That forced looking and left nothing behind: one browser_navigate satisfied it, and the regression test had to be recreated from memory afterwards. A committed spec is the verification AND the artefact — it re-runs on every later ticket, and it cannot be satisfied cheaply, because it has to pass against the running site. A session's Playwright MCP tools remain available for exploration, and the run still records the declared browser capability.

Custom gates (gates.custom) wire the repo's own commands — semgrep, behat, anything — as first-class gates: everything explicit (on, phase of code|test, single-line cmd), exit zero passes, and a command the shell cannot find reports tool missing, which blocks, never passes.

Three details worth knowing:

  • Anything that does not name a preset resolves to max — no file, an empty file, or a file that sets other things but never mentions one. A repo that has said nothing has not opted out of anything — tests-must-exist included, since 2.0. This is deliberately one rule rather than three: an earlier revision defaulted a file that exists to custom, which meant touch droost.workflow.yml turned mutation, playwright and coverage off and dropped PHPStan from max to 6, silently. If you want the gentler set, name it — preset: custom — so the choice is visible in a diff.
  • Thresholds never imply on. Writing coverage.min without coverage.on leaves the gate where the preset put it. An inferred switch would make min: 0 and on: false two spellings of one intent with two different failure modes. (phpstan.level: off was the one other switch; the 0.4 mandate superseded it — the attempt is noticed and the gate keeps the preset's level, so the recorded levers can never claim a gate was off when it ran.)
  • Something at the config path that is not a readable regular file is an error, not an absent config. A directory, a broken symlink or an unreadable file would otherwise swap your gates for the built-in ones and report nothing unusual.

Gate options

Most gates carry their thresholds inline — phpcs.standard, phpstan.level, coverage.min, mutation.msi_min. Every gate that spawns a tool also takes timeout (seconds) before the executor kills it and reports "could not run"; the default is 600, and xhigh/max give mutation 1800 and coverage 900 because infection over one kernel-test-heavy module already outruns ten minutes. One is easy to miss: rendered_check.routes is a comma-separated list of internal paths the live surface renders (routes: "/,/pricing"); omitted, it renders /. The render is an anonymous visitor's, so a page only some users may see is listed with the refusal it must give, /admin/reports@403, and a run declares one with declare-route /admin/reports --status=403; the gate then checks that it refuses, and a declared refusal that renders fails as public. On a Drupal site the render happens in a fresh drush process (droost:workflow:render-probe, answering in the gate result's own JSON), never inside the process running the gates: round 30 watched an in-process sub-request throw three times from an MCP server alive for hours while every fresh process rendered the same route fine, so the render leaves the process the way wiki_fresh always has. The option vocabulary is closed per gate — anything else is refused by name.

Every gate has a second switch after on: mode. block (the default, and the only behaviour that existed before) fails the phase on a blocking result. report runs the gate exactly as before but records the failure as REPORTED — the findings ride the report, the seeker reads them, the phase advances — and nothing reads as a pass. It is frozen into the run like on: whether a gate may block is not something a mid-run edit changes. The mandatory trio can never be put in report mode; that would be a disarm by another name, and the attempt is noticed and superseded like on: false.

Inherited debt — the adoption baseline

A legacy project fails these gates on debt the first change never touched: hundreds of phpstan errors, forty phpcs findings, config drift nobody will fix this month. The baseline is the operator's snapshot of that debt, written once from a terminal, committed beside the lever file in droost/baseline/:

drush droost:workflow:baseline --measure    # the bill: what each gate would inherit (read-only)
drush droost:workflow:baseline              # write it (operator terminal); commit droost/baseline/
drush droost:workflow:baseline --refresh    # re-measure: paid-off debt drops, growth is refused
drush droost:workflow:baseline --refresh --grow --reason="legacy import"   # accept more, on record
vendor/bin/droost-workflow baseline [--status|--measure|--refresh …]         # the same, standalone

From then on a consulting gate — phpcs, phpstan, eslint, stylelint, prettier, coverage, mutation, config_clean — answers two questions. Inherited is a finding the baseline records: reported, counted, never failing. New is everything else, and fails as before. The line reads passed — 0 new, 123 inherited, never a bare pass. A finding is keyed by file, rule, message and the text of its line, so a line that shifts stays inherited and a line that is edited becomes new — touching a file does not make its whole debt yours. phpstan runs through the baseline's own generated file; prettier's recorded files are inherited until the run touches them; coverage and mutation pass at or above their recorded floor when they miss the level's target, with the target named (the ratchet — the floor rises on refresh, never falls).

The run freezes the baseline's hash at begin. A baseline added, removed or edited under a run fails every consulting gate with "the baseline changed during the run"; the pack guard refuses the agent's shell from writing it and the agent's editor from touching the directory; the seeker names any move a defeat. baseline: { on: false } in the lever file is strict mode, one visible line. Design: docs/design-adoption-baseline.md.

Contributed gates

A module can contribute a gate the way it contributes a Drush command: a #[DroostGate] plugin (droost's droost_workflow submodule defines the type) declaring the command, the phases, a default mode and — required — a sentence saying what its verdict means. Enable the module and module:<id> joins every run at those phases, on by default, with the module recorded as provenance and the sentence repeated on a failure. droost_snyk is the reference: snyk test at code and test, report mode, until the site says otherwise:

gates:
  contributed:
    snyk: { mode: block }     # or { on: false }

Those two keys are all the lever file may set on a contributed gate — the command, phases and verdict are the module's contract, a different scan is a gates.custom entry. The dial never moves a contributed gate. Design: docs/design-contributed-gates.md.

Every door sees the same set. The drush commands and the MCP tools read the site's catalog directly; the standalone droost-workflow binary boots no Drupal, so it asks vendor/bin/drush droost:workflow:catalog before it resolves a lever. Status says which it was — levers.contributed_source — and when there is no drush, or the site cannot answer, it says the set was resolved from the lever file alone rather than presenting the shorter set as the whole (a run begun on that surface in the first live round was held to fewer gates than the same run through drush, and nothing said so). And a run that WAS begun blind — the binary run on a host where drush cannot reach the site — catches up: the first surface that can see the catalog weaves the missing gates into the run's record at the phase it is about to run, they run from then on (in the strict flow complete re-runs everything, so no run finishes without them; in the fast flow a late gate whose declared phases are all behind the run is woven into complete, so it still runs once), and the record says so — run.contributed_source names the door that began it, run.late_woven names each gate and the phase it joined at.

Work-item project config

The optional work_item: block is the project config of a tracker bridge (droost_jira and its kin). The engine never consumes it — it stays framework-free and knows nothing of Jira — except for the one provider it builds in, markdown (below, "Tickets (solo mode)"). It parses and validates the block so a typo surfaces in review rather than at the first write, and workflow:status echoes it. The shape is provider-agnostic on purpose: which tracker and cloud, which projects and issue types are workable, how branches are named, what the tracker calls its transitions, and a local map of every custom field under a name the site chooses, so no module ever hardcodes another team's field ids and a team's own layer only ever says developer_notes:

work_item:
  provider: jira
  cloud_id: <your Atlassian cloud id>
  projects: [PROJ]
  eligible_types: [Story, Task, Bug, Sub-Story]
  branch: { prefixes: { feature: feature, bugfix: bugfix }, base: development }
  transitions: { in_progress: 21, in_review: 121, done: 31 }   # for a /transition-style command; the engine never fires them
  fields:
    developer_notes: { id: customfield_10001, format: adf }
    testing_notes:   { id: customfield_10002, format: adf }
    developer_id:    { id: customfield_10003, format: user }
  track_map: { Bug: bugfix, Story: standard, Task: standard }
  writeback: { acceptance_criteria: description, dev_notes_field: developer_notes }
  status_map: {}          # empty is the common case: SCM events move the ticket, not droost
  publish: { target: confluence, space: DOCS, parent: "<page id>" }

A writeback target that names a fields entry resolves to its id. What a ticket MUST carry, and what goes into each field, is the bridge's or the team's business, never this block's — every write it describes stays gated.

Tickets (solo mode)

With droost and droost_workflow alone, tickets are markdown files in the repo, read through a WorkItemSourceInterface. work_item.provider: markdown turns it on; any other provider keeps the block's metadata-only meaning:

work_item:
  provider: markdown
  markdown: { dir: droost/tickets, prefix: TICKET }   # both optional; these are the defaults

A ticket is <dir>/open/<PREFIX>-<n>-<slug>.md, or under closed/ once it is done, with a YAML frontmatter holding at least title, status and ticket_number, and its sections after it. The file's name is its identity: its id is <PREFIX>-<n> and its number is n, and a ticket_number that says otherwise is kept as an extra key. Two older names read too: a split ticket, <PREFIX>-<n><letter>-<slug>.md (id <PREFIX>-<n><letter>, so a number can repeat), and a numberless <PREFIX>-<slug>.md (id the name, number null). A file named for the prefix in neither shape is refused by name:

---
title: TICKET-12-camp-news-by-recipe
status: ready
ticket_number: 12
type: feature
created: 2026-09-28
---

# TICKET-12-camp-news-by-recipe

## Summary

The states are backlog, ready, in_progress, review and done, machine names so a Drupal Workflow can use the same ids. The legacy open and closed read as ready and done, and a write always uses the new names. Every other frontmatter key is kept, and a move rewrites the status: line and nothing else, through a temporary file and a rename that keep the file's mode; a symlinked ticket or directory is refused.

  • run --ticket=<id> binds a run as it begins: status shows run.work_item, the work_item declaration is recorded as declare-changes --work-item records one, and the ticket moves to in_progress. When the run completes it moves to review, never to done, which is only ever a person's move. A move the source refuses is recorded as a note and never fails the run.
  • ticket list [--status=<state>], ticket show <id> and ticket new --title="…" [--type=<type>] are anyone's; new files into backlog, numbered one past the highest number any file's name carries, across open/ and closed/.
  • ticket move <id> <state> is the operator's (see the table below).

Tickets (cockpit mode)

work_item.provider: droost_cockpit reads tickets from a cockpit, a web service that owns a team's queue, and relays each run's events to it. The lever file names two environment variables and never holds either value, since the file is committed:

work_item:
  provider: droost_cockpit
  cockpit: { url_env: DRUPLIT_MAILBOX_URL, token_env: DRUPLIT_SEAT_TOKEN }   # path: /work-items/v1 is the default

url_env holds the cockpit's origin, and path is appended to it. token_env holds a bearer token, which is sent as Authorization: Bearer and never written, printed or put in an error. docs/cockpit-provider-api.md is the API a cockpit implements, and tests/fixtures/cockpit-stub.php is that API as a php -S router, held to the document by a test.

  • run --ticket=<id> fetches the ticket as the run begins and caches it in the state directory as work-item-<id>.json, so a restart with the cockpit down still binds it. A ticket the cockpit has never answered for is refused.
  • ticket list, ticket show and ticket new ask the cockpit. ticket move exits 2 with "move tickets in the cockpit": its queue owns ticket state, and it reads a run's progress from the run's events.
  • Every event the log gains is posted as it is written, and once more at the end of every command, from the cursor in relay-cursor, in batches of at most 100 events and 256 KiB. Delivery is at least once, and the cockpit dedups by event_id. The first failure in a process is reported once on stderr and silences the relay for the rest of that process, so a cockpit that is down costs a run at most one 5-second timeout and never fails it. The events wait in the log.
  • relay sends what waits and exits 0 when nothing does, or 1 while events still wait. status shows relay: {cursor, pending, last_error}.
  • Requests go over http or https only, with redirects refused rather than followed, since a redirect would carry the token to another host.

Unknown keys are errors

A loader that shrugs at phpstain: hands back a run with static analysis quietly disabled and a report that says everything passed. So every unknown setting, gate, option, phase, mode and preset is refused by name:

droost.workflow.yml: unknown gate "phpstain" (known: phpcs, phpstan, eslint,
stylelint, prettier, phpunit, mutation, playwright, parity, coverage,
rendered_check, config_clean, grounding_check, composition_check, wiki_fresh)

Which gates run when

WHETHER a gate runs is the lever file's business. WHEN it runs is the engine's phase map, frozen into each run when it begins:

plan: none
code: phpcs, phpstan, eslint, stylelint, prettier, config_clean, grounding_check, composition_check
test: phpcs, phpstan, eslint, prettier, phpunit, mutation, playwright, parity, coverage, rendered_check, config_clean
complete: phpcs, phpstan, eslint, stylelint, prettier, phpunit, mutation, playwright, parity, coverage, rendered_check, config_clean, grounding_check, composition_check, wiki_fresh

Plan runs no gate — there is no code yet to measure. droost's Drupal half adds one check there, plan_consulted: plan waits until the spec as it stands was put to droost (droost_consult). That is the one thing about how a site is built that the workflow forces (owner, 2026-10-02: "we force the workflow … but droost is the question and answer"). What is built after it is the agent's choice: grounding_check and composition_check record it beside droost's advice, a build rule in rules: is advice (block is read as report), and only a record that is false fails, such as a page declared as one owner and served by another. Code gates the diff with static analysis, and config_clean asks the booted site whether a fresh drush config:export would change the tracked tree: Drupal compares configuration as data, so a hand-written file that imports fine can still diverge from the canonical serialization, and the divergence surfaces later as a noisy re-export burying real changes. Zero diff, against a target git can see, is the passing state. Test runs the functional gates, and re-runs the static ones over the tests it has just written — a suite whose own source would not pass code's standards is not a suite whose green anyone should believe. Complete opens by capturing what was built — the documentation work that was its own phase until 0.4 — and then re-runs the full enabled set as the terminal safety net, custom gates included, so every other enabled gate is met at least twice: once at its own phase, once at the end.

That is the strict flow, which high and the levels above it run. Since 0.11 low and medium run the fast flow (flow: fast | strict, a lever the level defaults): each gate runs once, at the phase that owns it. Code runs the analysers and the unit tests (phpunit, mutation and coverage where on), test runs playwright, parity, the rendered check and config_clean, and complete runs no gate but wiki_fresh on the documentation it writes. A run freezes its flow's table as phase_gates when it begins. The same two levels default scope: ticket on phpcs, phpstan, phpunit and playwright: the analysers read the files the run changed, and the suites run the test files it added or changed, the Playwright CLI handed those specs. scope: full runs a gate's whole set. A scoped row says so ("ticket scope: 2 file(s) this run changed, not the full suite"), and a required suite the ticket gave no test to run fails. At low the browser gate also takes reuse: recorded: droost-workflow specs, run at test, runs the ticket's specs exactly as the gate would and records the result against a fingerprint of the tree, and the test phase then takes that record rather than running the suite a second time. The row says so ("recorded at …, on this tree: not run again"). An edit after the recording, a tool that could not run, or a level at reuse: never (every level above low) runs the suite. What is taken is always droost's own run of the tool, never an account of one. P6 run 23 measured what the strict flow costs at the bottom of the dial: its 279-test browser suite ran three times in 86 minutes, for a ticket whose own tests were two files. wiki_fresh is due only here, and only here CAN it be true: it asks the site whether the project's own documentation still matches the code, and complete is the phase that just wrote it. A stale page is read as fact, which is worse than no page.

Promise against proof: the criteria table

The full spec's acceptance-criteria table carries a Verified By column, empty at plan. The test phase fills it with the test that proves each row — the PHPUnit method or class, or the Playwright spec — or manual — <reason> for a criterion no test can prove. At complete an empty cell (or a missing column) is recorded, not blocking: the report names each unverified row under criteria_verified, prints manual as manual, never as passed, and workflow:status shows the three lists. It was a refusal once, and the shape of a markdown file ended three live runs whose every gate had passed (F-35), so the spec's shape is a record now and the gates are the wall. A criterion verified through verify_criterion is recorded against the test that ran. A quasi-spec at medium/low has no table and is not held to one.

The feedback loop

A blocking gate does not end a run; it starts a bounded loop. Each failing invocation spends one attempt per blocking gate — recorded in run state as feedback_attempts, measured against max_gate_retries — and the agent fixes the cause between invocations. max_gate_retries: 2 means one attempt plus two retries; 0 means one attempt and no retry. A missing tool spends budget exactly like a failure, because a missing binary re-invoked forever is the worst infinite loop of all.

When the budget is spent, the phase is recorded failed — terminal — and run refuses to execute anything further. Every surface renders the same envelope: {outcome, current_phase, report, awaiting, retries, loop}, where retries.exhausted separates "fix it and run again" from "this run is over". Exit codes stay simple — paused is not failed, and both kinds of failure exit non-zero. Recovery from a terminal failure is deliberate: vendor/bin/droost-workflow reset (or drush droost:workflow:reset) archives the record to droost/droost-workflow/history/ and clears the way — the same verb that closes out a COMPLETED run, whose record also persists until reset.

The fast flow's loop: back to code, or a follow-up

In the fast flow a failure at test does not retry in place: the run goes back to code (outcome: returned, exit non-zero), with the failing gates named in blocked. Code's gates run again on the next run, and test after them. max_loops bounds it (low 2, medium 3, high and xhigh 5, max 5). A return because the code moved spends nothing: when the files code's gates measured change after they ran (a fix made at test, an edit at complete), the next test or complete sends the run back to code before any gate runs, so no green stands over code it never saw. Markdown and the workflow's own droost/ directory are not code a later phase re-measures.

A failure that is out of the ticket's scope (a full-scope browser suite whose every failure is in a spec the run did not change), or one still there when the budget is spent, becomes a follow-up ticket: the phase is recorded deferred, never passed, a follow_up row at the phase names each ticket, and the run moves on (outcome: deferred, exit 0). follow_ups decides where they go: auto (the default) writes to the cockpit when it is the project's ticket source and to markdown otherwise (the project's own ticket directory, or docs/tickets/); markdown and cockpit say so outright; none writes nowhere, so a spent budget fails the phase. At max a spent budget fails the phase too. A follow-up the cockpit refuses is written as markdown, and the record says why. A tool that could not run (missing, crashed, timed out) is not code to fix, and keeps the in-place retry above. Every return, its reason and each follow-up are in the envelope's loop block and in run.json.

Run state

Run state lives beside the lever file, in droost/droost-workflow/run.json (a project that still has the older .droost-workflow/ directory keeps using it until that directory is gone):

{ "v": 1, "run_id": "...", "phases": { "plan": "passed", "code": "active" } }

On the filesystem for the same reason the levers are. If run state lived only in Drupal's State API, a run started against a live site could not be resumed, inspected, or even described from a plain checkout — and the two surfaces would be two pipelines sharing a name.

Writes go through a temporary file and a rename, so a run that dies mid-write leaves the previous state intact. A state file that cannot be parsed is never deleted or replaced: it is still evidence of what a run was doing.

Two limits on that, stated because an overstated durability promise is worse than a modest one: there is no fsync, so a process crash is covered but a power cut is not; and there is no locking, so two processes doing load-modify-save against one file will lose an update without either being told. Nothing is ever torn. The model is one run per repo.

A phase that failed is never quietly recorded as passed. Advancing stamps the phase you are leaving as passed, so advancing away from a failed or skipped phase is refused outright — clearing a failure has to be a deliberate act, not a side effect of moving on. Advancing backward, or advancing a run that has already reached its terminal gate, is refused for the same reason: a report has to be able to describe the run honestly. The fast flow's one way back is its own transition, test or complete to code, which sets the phases between back to pending; and a phase it leaves with follow-ups is deferred, a word of its own, never passed.

Run events

Beside run.json, the engine keeps events.jsonl: an append-only log of every state change it persists, one JSON event per line, written after the save. It is the record a relay or a dashboard reads (droost-workflow events [--after=<seq>] prints it), and it works with nothing else installed:

{"schema":"droost.run-event/1","event_id":"evt-4f1c2a9be0d37a61","seq":7,"run_id":"run-8c2d4e6f1a3b","work_item_id":"TICKET-12","at":"2026-09-28T15:04:11+00:00","type":"phase.attempted","payload":{"phase":"code","attempt":1,"report":{"phase":"code","advance":false,"tally":{"passed":3,"failed":1},"gates":[]}}}
Type When Payload
run.started a run begins preset, mode, enforcement, phases, spec, base_commit
phase.began a phase becomes current phase
phase.attempted a phase's gates ran, pass or fail (a failed gate never ends a phase, so this is the only event that shows one) phase, attempt, report (the phase report: advance, tally, every gate)
phase.ended the run leaves a phase phase
question.asked a save leaves a question newly awaiting question_id, question
question.answered answer is saved question_id, answer
run.completed the last phase ended outcome (the run envelope)
run.reset reset archived the run archived_run_id
phase.returned the fast flow sent the run back to code from, to, reason (failed or moved), spent, max_loops, gates
follow_up.filed a deferred phase's failure was written up phase, gate, why (spent or outside), summary, title, id, source, path, at, note

Every event has schema, event_id (evt- and 16 hex digits, the consumer's dedup key: delivery is at least once), seq (strictly increasing per log from 1, a relay's cursor and never an identity), run_id, work_item_id (the bound ticket, or null), at, type and payload. schema/run-event.v1.json is the JSON Schema. Version 1 is additive only: new types and payload keys may come, a consumer ignores what it does not know, and renaming or removing one is droost.run-event/2.

The log spans runs, so reset keeps it. An append holds an exclusive lock, so two processes never interleave a line or reuse a seq (on a state directory over NFS that lock is advisory, as flock is). A log that cannot be written is reported once on stderr and never fails or blocks the run. It lives in the state directory, which init keeps out of version control, as it does evidence.sqlite.

The pack

The phases ship as a .claude/ pack — seven skills (the three entry verbs workflow-start, workflow-continue and workflow-status, and one skill per phase), three slash commands that are one-paragraph pointers to the verb skills (/droost:workflow:start, /droost:workflow:continue, /droost:workflow:status — Claude Code's way in; every other host reads the skill itself, which is why the procedure lives there and not in the command), five agents (the plan researcher and spec-writer, the adversarial workflow-seeker, the one-finding-at-a-time workflow-bug-fixer, and droost-debugger, which flips xdebug on for a stubborn failure and back off after), and a shared partial on using droost. Installing it into a repo writes:

.claude/skills/workflow-{start,continue,status}/SKILL.md
.claude/skills/workflow-{plan,code,test,complete}/SKILL.md
.claude/commands/droost/workflow/{start,continue,status}.md
.claude/agents/workflow-{researcher,spec-writer,seeker,bug-fixer}.md
.claude/agents/droost-debugger.md
.claude/hooks/droost-workflow-guard.php
.claude/partials/droost-usage.md
.claude/templates/evaluation.md
droost.workflow.yml          # only if you don't already have one

Each phase skill states four things: its entry gate, the work, its exit gate, and what it can and cannot check without a booted site. That last section is the point. Nearly every droost tool needs a running site, so a CLI run has real blind spots — and a run that hides them produces a report nobody should trust.

The guard hook is the one host-specific piece — the table under Install says what a host without it loses: it is a Claude Code pre-tool hook, and it refuses the operator's commands from the agent's shell, plus any agent edit under droost/baseline/, under droost/droost-workflow/, or to the guard itself. There are five, and each is a loosening somebody has to own:

Each verb has TWO spellings — the drush one and the standalone binary's — and the guard refuses both. It has to: droost-workflow bypass "x" and drush droost:workflow:bypass "x" are the same act, and a wall that knows only one of them is not a wall. The refusal names the spelling that was used, so the command it hands the operator is one they can actually run.

Refused from the agent's shell What is still the agent's to run
droost:workflow:gate-waive, droost-workflow gate-waive — nothing; a waiver is always the operator's
droost:workflow:baseline, droost-workflow baseline (writing) --status and --measure, which only read
droost:workflow:bypass, droost-workflow bypass bypass --off, which tightens
droost:workflow:effort <level>, droost-workflow effort <level> bare effort (reports) and effort <level> --preview (prices it)
arming a write gate — droost:gate allow_* on, or the config:set droost.settings allow_* true form disarming the same gate, which tightens
droost-workflow ticket move <id> <state> ticket list, ticket show and ticket new (a follow-up, filed in backlog)

The pattern is the same in each row: the reading and the tightening are the agent's, the loosening is not. An agent that needs one proposes it — the refusal prints the exact command to hand over, and in Claude Code the operator runs it with ! drush … (or ! droost-workflow …) so it lands in the same transcript. On a host without pre-tool hooks the status document's run half says so — enforcement.effective: advisory — because the gates still hold the run server-side but nothing stops an out-of-phase edit, and a report must not claim a discipline the host never had.

Ownership is explicit. Every directory the pack owns gets a .droost-workflow-pack marker. Re-running the installer refreshes those directories and nothing else; a directory without the marker belongs to you and is refused rather than overwritten. Your droost.workflow.yml is never refreshed at all — it is version-controlled intent you wrote, and resetting your gates on an unrelated re-install would be an unpleasant surprise.

Two files outside the pack are touched, and only additively: the hook entries are merged into .claude/settings.json (an existing file keeps everything else it holds; one that cannot be parsed is refused, not replaced), and the run-state directory's ignore lines are appended to .gitignore if they are not already there. Both are edited in place rather than announced in init's file tally, so git diff after an install is where you see them.

Install

This package is a framework-free PHP library. It requires no Drupal, and it is the whole standalone surface:

composer require --dev droost/workflow
vendor/bin/droost-workflow init      # writes the pack + a default lever file

What init writes, and for which host

Be precise about this, because "host-agnostic" is easy to claim and was overstated here until 2026-09-15. Two of the three surfaces are genuinely host-neutral; the third is Claude Code's, and there is no second implementation of it.

What Who reads it
AGENTS.md — the pipeline's own block: that a run comes before a write, and that the discipline is enforced every host. The cross-tool convention. init CREATES the file when a project has none, because an installed pipeline no agent is told about is the failure this block exists to prevent
droost.workflow.yml — the levers every host, and a person. A file rather than site config partly so a reader with no site can see what a run is held to
vendor/bin/droost-workflow — the verbs, including the operator's three every host. Anything that can run a command can drive a run
.claude/skills/workflow-* — the phase procedures Claude Code loads them by name; any host can read them as prose, and they are written to be read that way, which is why the procedure lives in the skill and not in the slash command
.claude/commands/droost/workflow/* — three one-paragraph pointers Claude Code only
.claude/hooks/droost-workflow-guard.php + its settings.json wiring Claude Code only. This is the enforcement, and it has no equivalent anywhere else

So a Codex or "other" user gets the doctrine, the levers, the CLI and the skills as readable prose — and not the enforcement. declare-tasks codex|other|none records which host is driving, and on a host with no pre-tool hook the status document says enforcement.effective: advisory in as many words: the gates still hold the run, because the engine runs them, but nothing stops an out-of-phase edit. A report must not claim a discipline the host never had.

Porting the guard is the only work that would change that row, and nobody has done it. Said plainly here so the gap is a known limitation rather than something a reader discovers by reading PackManifest.

The Drupal surface ships with droost, not here. drupal/droost's droost_workflow submodule supplies the two things that genuinely need a booted site — the drush droost:workflow:* commands and the droost_workflow_status / droost_workflow_run MCP tools. Enable it and you get both; the pipeline underneath is this library either way.

That split is P6.7 in droost's roadmap, and the reason is delivery rather than capability: a Drupal site builder does not install contrib from a git remote, and a plain Claude Code or Codex user does not want an AI work pipeline delivered as a Drupal module. Nothing about the pipeline changed.

The lever file's preset is a scalar (preset: custom) — there is no presets: block to configure; a preset is a base the gates: entries overlay.

Two surfaces, one pipeline

The same run, the same levers, the same report — whether or not there is a site.

# Standalone. Any Drupal repo, no booted site, nothing running.
vendor/bin/droost-workflow init
vendor/bin/droost-workflow status
vendor/bin/droost-workflow run
vendor/bin/droost-workflow reset      # archive a finished run, start clean
vendor/bin/droost-workflow baseline --measure   # the inherited-debt bill; a bare `baseline` writes it

# Against a live site.
drush droost:workflow:status
drush droost:workflow:run
drush droost:workflow:answer "yes, continue"
drush droost:workflow:swap agentic
drush droost:workflow:reset
drush droost:workflow:baseline --measure        # the same bill, config_clean included

The only thing that differs is what the site-dependent gates can say. Verified on a real Drupal 11.4.4 site:

Surface rendered_check
live (drush) passed — "1 route(s) rendered", a real render of / in a fresh drush process
standalone skipped, no site — with the reason recorded

That difference is the entire point. The CLI surface is not a degraded live run pretending otherwise; it is a run that tells you exactly which checks it could not perform. Everything else — which gates ran, their verdicts, the phase, the advance decision — is identical, because both surfaces call one facade and differ only in which site driver they inject.

The MCP surface (ships with droost)

drupal/droost's droost_workflow submodule exposes the same engine over MCP, as a third and fourth front onto the one WorkflowFacade:

  • droost_workflow_status — read-only. The resolved levers (which preset, which gates, and where that came from), the phase order with each phase's status, the latest tally, and whether the run is awaiting an answer.
  • droost_workflow_run — drives the run: gates the current phase and advances when they pass. answer and swap are ARGUMENTS of this one tool rather than separate tools, because they are sub-operations of driving a run and every extra destructive tool is another separately allow-listed surface.

Both take an optional project (absolute path to the repository); omit it for the site's own root. A root that is not a directory comes back as a failure envelope naming the path — never an exception, which over JSON-RPC would tell the caller nothing it could act on.

It is a SUBMODULE so a plain consumer of this package never pulls the alpha mcp_server: enable it only if you want the MCP surface. It depends on droost:droost, mcp_server:mcp_server and this module.

droost_workflow_run is STDIO/Drush-only — gated on the transport alone, no allow_* flag, the same posture droost_verify has for the same risk class (it spawns the project's own analysis binaries). Its whole body runs inside droost's Fiber shield, because a run can reach the rendered check and Drupal's renderer suspends the fiber in a way the MCP SDK misreads as a dropped response.

Analysing it. The submodule lives in the drupal/droost repository, not here, and is type-checked there by that module's own gates (drupal.org GitLab CI at PHPStan level max). This package's ./scripts/lint covers the engine and the pack only.

Requirements

PHP 8.3+. The engine's only runtime dependency is symfony/yaml — no Drupal bootstrap is required to read config or run state, which is what lets the same code serve both surfaces.

Development

composer install
scripts/lint            # phpcs, phpstan (level max), phpunit
scripts/lint src/Config # scope phpcs to a subset

All three legs are hard. A missing binary fails the gate rather than skipping it and says which invocation it could not run — an environment that cannot run a gate is a broken environment, not a gate that does not apply.

License

GPL-2.0-or-later.