milpa / app-runtime
The agent runtime a Milpa app INSTALLS instead of copying: session gate, sub-agent delegation, tree budget, sterile-loop guard and the live bridge. Lives here so an existing app receives its improvements — a template that copies what nobody edits is a package in disguise.
Requires
- php: ^8.3
- milpa/command: >=0.25.2 <1.0
- milpa/console: >=0.20 <1.0
- milpa/container: >=0.1 <1.0
- milpa/core: >=0.12 <1.0
- milpa/http: >=0.1.5 <1.0
- milpa/live: >=0.25 <1.0
- milpa/live-tui: >=0.7 <1.0
- milpa/live-web: >=0.29 <1.0
- milpa/plugin: >=0.17.1 <1.0
- milpa/runtime: >=0.15 <1.0
- milpa/tool-runtime: >=0.17.1 <1.0
- nyholm/psr7: ^1.8
Requires (Dev)
- friendsofphp/php-cs-fixer: ^3.75
- milpa/agent: >=0.44 <1.0
- milpa/ai-gateway: >=0.25 <1.0
- milpa/auth: >=0.9 <1.0
- milpa/data: >=0.2.2 <1.0
- milpa/devtools: >=0.26 <1.0
- milpa/event-store: >=0.3 <1.0
- milpa/events: >=0.4 <1.0
- milpa/mercure: *
- phpstan/phpdoc-parser: ^2.3
- phpstan/phpstan: ^2.1
- phpunit/phpunit: ^11.5
Suggests
- milpa/agent: The agent runtime itself: sessions, gates, sub-agents. Without it the agent operations are not offered.
- milpa/ai-gateway: One message and tool-call shape across model providers. Needed by the agent operations.
- milpa/auth: Verifying the token a caller presents over HTTP.
- milpa/data: Persisting those tokens.
- milpa/devtools: `coa doctor`, `coa repair` and `coa update`.
- milpa/event-store: Append-only session streams. Needed by the agent operations.
- milpa/live-web: Mounts the live wire (LivePlugin): server-rendered components take actions over HTTP with signed state, CSRF and replay protection — dashboards, admin panels.
Provides
None
Conflicts
- milpa/agent: <0.44
- milpa/ai-gateway: <0.25.0
- milpa/command: <0.25
- milpa/event-store: <0.3
Replaces
None
- dev-main
- v0.161.0
- v0.160.2
- v0.160.1
- v0.160.0
- v0.159.1
- v0.159.0
- v0.158.4
- v0.158.3
- v0.158.2
- v0.158.1
- v0.158.0
- v0.157.7
- v0.157.6
- v0.157.5
- v0.157.4
- v0.157.3
- v0.157.2
- v0.157.1
- v0.157.0
- v0.156.1
- v0.156.0
- v0.155.0
- v0.154.1
- v0.154.0
- v0.152.0
- v0.151.4
- v0.151.3
- v0.151.2
- v0.151.1
- v0.151.0
- v0.150.1
- v0.150.0
- v0.149.2
- v0.149.1
- v0.149.0
- v0.148.0
- v0.147.1
- v0.147.0
- v0.146.0
- v0.145.2
- v0.145.1
- v0.145.0
- v0.144.2
- v0.144.1
- v0.144.0
- v0.143.0
- v0.142.0
- v0.141.0
- v0.140.0
- v0.139.1
- v0.139.0
- v0.138.0
- v0.137.0
- v0.136.1
- v0.136.0
- v0.135.0
- v0.134.0
- v0.133.3
- v0.133.2
- v0.133.1
- v0.133.0
- v0.132.1
- v0.132.0
- v0.131.1
- v0.131.0
- v0.130.2
- v0.130.1
- v0.130.0
- v0.129.0
- v0.128.1
- v0.128.0
- v0.127.0
- v0.126.0
- v0.125.1
- v0.125.0
- v0.124.0
- v0.123.1
- v0.123.0
- v0.122.0
- v0.121.0
- v0.120.2
- v0.120.1
- v0.120.0
- v0.119.0
- v0.118.0
- v0.117.0
- v0.116.0
- v0.115.0
- v0.114.0
- v0.113.0
- v0.112.1
- v0.112.0
- v0.111.0
- v0.110.0
- v0.109.0
- v0.108.0
- v0.107.0
- v0.106.0
- v0.105.0
- v0.104.0
- v0.103.0
- v0.102.1
- v0.102.0
- v0.101.0
- v0.100.0
- v0.99.0
- v0.98.0
- v0.97.0
- v0.96.0
- v0.95.0
- v0.94.1
- v0.94.0
- v0.93.0
- v0.92.0
- v0.91.1
- v0.91.0
- v0.90.0
- v0.89.0
- v0.88.0
- v0.87.0
- v0.86.1
- v0.86.0
- v0.85.0
- v0.84.0
- v0.83.0
- v0.82.0
- v0.81.0
- v0.80.0
- v0.79.0
- v0.78.0
- v0.77.0
- v0.76.0
- v0.75.0
- v0.74.0
- v0.73.1
- v0.73.0
- v0.72.0
- v0.71.0
- v0.70.0
- v0.69.1
- v0.69.0
- v0.68.0
- v0.67.0
- v0.66.0
- v0.65.0
- v0.64.1
- v0.64.0
- v0.63.0
- v0.62.0
- v0.61.0
- v0.60.0
- v0.59.1
- v0.59.0
- v0.58.0
- v0.57.0
- v0.56.0
- v0.55.0
- v0.54.0
- v0.53.0
- v0.52.0
- v0.51.0
- v0.50.0
- v0.49.1
- v0.49.0
- v0.48.3
- v0.48.2
- v0.48.1
- v0.48.0
- v0.47.0
- v0.46.0
- v0.45.0
- v0.44.1
- v0.44.0
- v0.43.1
- v0.43.0
- v0.42.1
- v0.42.0
- v0.41.1
- v0.41.0
- v0.40.0
- v0.39.0
- v0.38.0
- v0.37.0
- v0.36.0
- v0.35.0
- v0.34.0
- v0.33.2
- v0.33.1
- v0.33.0
- v0.32.1
- v0.32.0
- v0.31.0
- v0.30.0
- v0.29.0
- v0.28.1
- v0.28.0
- v0.27.1
- v0.27.0
- v0.26.0
- v0.25.0
- v0.24.0
- v0.23.1
- v0.23.0
- v0.22.1
- v0.22.0
- v0.21.0
- v0.20.0
- v0.19.0
- v0.18.0
- v0.17.1
- v0.17.0
- v0.16.1
- v0.16.0
- v0.15.2
- v0.15.1
- v0.15.0
- v0.14.0
- v0.13.0
- v0.12.4
- v0.12.3
- v0.12.2
- v0.12.1
- v0.12.0
- v0.11.1
- v0.11.0
- v0.10.0
- v0.9.0
- v0.8.0
- v0.7.0
- v0.6.0
- v0.5.0
- v0.4.1
- v0.4.0
- v0.3.0
- v0.2.0
- dev-rodrigoteamx/delivery-final
- dev-rodrigoteamx/termination-consumer
- dev-rodrigoteamx/delivery-closure
- dev-rodrigoteamx/review-selection
- dev-rodrigoteamx/acceptance-sdk
- dev-rodrigoteamx/trial-verdict
- dev-rodrigoteamx/candidate-sdk
- dev-rodrigoteamx/enforce-withdrawal
- dev-rodrigoteamx/resident-observer
- dev-rodrigoteamx/witnessed-failures
- dev-rodrigoteamx/trial-freshness
- dev-rodrigoteamx/effect-evidence
- dev-rodrigoteamx/bounded-recovery
- dev-rodrigoteamx/trial-drain
- dev-rodrigoteamx/recovery-offer
- dev-rodrigoteamx/progress-recovery
- dev-rodrigoteamx/delivery-verdict
- dev-rodrigoteamx/component-parts
- dev-rodrigoteamx/agent-ui-authoring
- dev-rodrigoteamx/screen-drafts
- dev-rodrigoteamx/todo-resources
- dev-rodrigoteamx/screen-registry
- dev-rodrigoteamx/screen-preview
- dev-rodrigoteamx/identity-panel
- dev-rodrigoteamx/plugin-authority
- dev-feat/the-judge-gets-someone-to-judge
- dev-fix/the-ceremony-derives-its-urls
- dev-fix/every-taught-command-runs
- dev-feat/a-capability-declares-its-plugin
- dev-feat/the-reader-learns-to-refuse
- dev-feat/window-says-both-numbers
- dev-feat/capabilities-enable-http
- dev-feat/agent-real-token-cost
- dev-fix/passkey-plugin-metadata
- dev-feat/enroll-prefers-cross-platform-authenticator
- dev-slice/evidence-receipt
- dev-feat/evidence-predicate
- dev-fix/dry-run-readonly-ceiling
- dev-feat/intra-leg-wiring
- dev-feat/progress-wiring
- dev-feat/house-context
- dev-feat/operation-contract
- dev-feat/work-protocol-graduation
- dev-feat/debt-signals
- dev-feat/intent-claim-admissibility
- dev-feat/window-budget
- dev-feat/scoped-grants-and-closure
- dev-feat/reasoning-wiring
- dev-feat/lazy-toolbox-wiring
- dev-fix/window-aware-compaction-factory
- dev-feat/board-asset-base-config
- dev-fix/agent-optional-again
- dev-feat/register-state-machine
- dev-feat/live-render-path
- dev-feat/live-render-helper
- dev-feat/recipe-apply
- dev-feat/operations-declare-their-subject
- dev-fix/reach-command-0-7
- dev-feat/foundation-as-a-governed-act
This package is auto-updated.
Last update: 2026-09-15 02:26:26 UTC
README
milpa/app-runtime
The agent runtime a Milpa app installs instead of copying.
What an agent is allowed to do inside your app, what your app knows how to do, and the two surfaces you drive it from — the CLI and the agent screen. All of it arrives by version.
Declared screen pages
With LivePlugin enabled and live.secret configured, GET /live/page?component=<name> returns a
complete HTML document. It loads the local runtime, remote runtime and Alpine once, plus the shipped
Milpa design styles, local fonts, and every rendered descendant's declared styles, scripts and messages.
The document works on its own or inside the panel's preview iframe. live.route changes the page,
endpoint and design-asset mount; runtime URLs retain their existing root mounts.
screen:declare validates the entire props.children tree before writing. An unknown or malformed
child returns ok: false with its path and preserves the previous screen; no served-evidence receipt
is issued. Invalid trees already in the store return HTTP 422 instead of rendering a partial screen.
screen:types reads the current component and renderer registries. It reports canonical contract
names with an HTML renderer and explains unavailable registrations. The declaration's type field
references that operation with x-milpa-source; validation reads the registry when called, including
plugins that boot after LivePlugin.
A plugin can extend the live door with an object that already has its collaborators:
// Run after LivePlugin boots. These are the same registries used by GET and action POST. $container->get(\Milpa\Live\Contracts\Component\ComponentRegistryInterface::class) ->register('task-item', new TaskItem($repository)); $container->get(\Milpa\Live\Rendering\ComponentRendererRegistry::class) ->registerFor('task-item', new TaskItemRenderer($container->get( \Milpa\Live\Contracts\Transport\StateTransferCodecInterface::class, )));
The registration name must match TaskItem::contract()->name. Its renderer supplies the component
HTML and signed state envelope. With live-web 0.29+, x-data="milpaComponent({componentId: 'task'})"
and @click="act('toggle', {})" use the existing signed transport and HTML reconciliation. No app
transport module is required. A declared screen can use this type at its root or inside a supported
container's props.children. Renderer DeclaresClientAssets and component presentation resources
are collected from descendants. Merely advertising a class through DeclaresComponents does not
construct it or register a renderer. An opaque registry must implement ListsComponents to offer
types through discovery. Existing configured components and built-in types remain supported.
This page shows the current declaration. It is not an isolated draft or a deployment boundary, and rendering a group of controls does not yet provide shared application state between them.
Why this package exists
Because it used to live inside the template, and that meant it never reached anyone.
milpa/framework is type: project. When you run composer create-project, its src/ is copied
into your app and from that moment it is yours. That is exactly right for the example plugin you are
going to delete. It is exactly wrong for the agent runtime, which improves every week and which
nobody ever edits.
The symptom that exposed it, measured: an app created one day earlier did not receive the
permission-question buttons, or the indicator that pulses on every real event, or agent:board —
even after updating everything. And the worst case was the quiet one: it did receive the new
milpa/live-tui, which knows how to paint what the system said in a different colour from what the
model said, and saw no change at all — because its copied screen never emitted the markers that
trigger that painting. Half the improvement landed, half didn't, and nothing said so.
The rule that came out of it, and that this package applies: you copy what you are going to edit; you install what you are going to use. A template that copies files nobody will touch is a package in disguise — all of a package's cost, none of its benefit.
What's in it
The gates — what an agent may do
| piece | what it decides |
|---|---|
SessionToolGate |
whether a call proceeds: permission, intent contract, sterile loop, ordering |
SubAgentSpawner |
delegating to a child session and resuming it — with fresh context, not re-delegating |
TreeBudget |
how many steps the tree spends, not each child: bounding the child does not bound the tree |
SterileLoopGuard |
not repeating a call that already failed the same way twice — on by default: at its home tolerance it would have refused 81 of a sick run's 89 calls and none of the healthy runs'. Opt out with agent.sterileLoopGuard: false; an integer sets the tolerance |
PrerequisiteGate |
an ordering obligation, executed: until the required thing runs, the rest does not. The system renews a session's standing obligation with a cheap read of its own state (agent_show) — orientation, not curation: a turn opened by bookkeeping becomes a bookkeeping turn (measured, twice). agent.renewalTool names another tool; false disables renewal, declared — never silent |
SessionOptionTable |
withdrawing a tool from a session's catalogue — forbidding, not asking |
BroadcastingEventStore · SurfaceBroadcaster · MercureBroadcaster |
getting what happens to the live surfaces while it happens |
SessionBookkeeping · SessionPlanBoard |
the session's plan and to-dos, bound to its id |
The resident resolves an optional TrialInputObserver registered in the app's DI container:
$container->registerService(TrialInputObserver::class, $hostObserver);
Register it during trusted host composition, before invoking the agent. AgentOperations passes
that same object to its shared trial runner; the existing trial event reports whether an attempt
has a known, partial or unknown witness. A wrong service type or a resolution error is a visible
configuration failure, never silently replaced by an unobserved runner. Disabled trials do not
resolve the observer. A hook or capture failure preserves the trial verdict and records unknown.
No observer binding keeps the previous behavior. The host still supplies and operates its capture
mechanism; registering a hook does not install a tracer or attest complete inputs. Measured in
Greenhouse decisions/0351 and evidence/0668.
Hosts that compose a TrialRunner directly can also pass inputObserver.
Its before(TrialInputAttempt) and after(TrialInputAttempt, int $exit) hooks surround the native
test execution. The observer runs outside the tested process; tool output is never this channel.
The returned record binds id, copy, operation, and arguments to that attempt and declares
scope: copied-app-file-content-presence-and-directory-members/v1, complete_execution_inputs: false, status: known|partial|unknown, and inputs. Each relative input carries facets
(content, presence, or members) and its before state (kind, plus sha256 for file content
or members for enumerated directories). The observer must report writes during execution,
incomplete resolution, and missing captures conservatively; a post-execution hash alone cannot
attest an input. The runner rejects records belonging to another attempt.
When the gate and executor share their TrialRouter and session, the runner's witness reaches
SterileLoopGuard through a one-use host channel. Known changed inputs permit new work while
unrelated edits retain the old failures. Returning to old inputs restores their failure history;
success clears only the exact known input identity. Missing or partial observations cannot prove
a repair, and an unreadable current input retains its failures. The native trial event records
the attempt, scope, status and identity. Without an observer, the existing argument-based behavior
is unchanged. No tracer, platform dependency, or observer is enabled by default. This bounded
file scope excludes vendor, var, cache, .env, the trial runner, external paths, environment,
clock, randomness and services; it does not claim complete or semantic dependencies. Measured in
Greenhouse decisions/0350 and evidence/0667.
During progress recovery, ConsentBridge removes declared reads from the offered catalogue using
the session gate's current state. Successful material work, recorded evidence or a completed todo
restores them; failed writes and pending confirmations do not. Durable option removals remain in
force. Full and lazy discovery use this current offer, including previously discovered schemas.
Offering a mutation does not authorize it: scopes and argument-dependent effects are still judged
when it is called. Catalogue inspection does not execute that judgment or open consent questions.
SessionProgressProbe opens recovery after four model calls without recorded growth. It allows
one further window of the same size for preparation, then reports exhaustion if growth is still
absent. A successful artifact-producing operation, recorded evidence or a completed todo resets
the window, including on its last call. Plan edits, repeated todos, failures and confirmation
requests do not. Observations explicitly distinguish pending, recovered and exhausted recovery;
an unavailable store cannot claim any of them. The orchestrator enforces this contract without
changing tool permissions or the total step budget. Windows belong to the current invocation.
The operations — what your app knows how to do
AgentOperations, SessionOperations, CapabilityOperations and TokenOperations are the operation
groups a Milpa app registers. They are returned, never self-registered: whoever assembles the
registry decides which groups get in and with what authority, and a group that registered itself
would take that decision away.
At a natural end, closure.verified covers scope: recorded_work: it requires positive
recorded evidence, no open or unevidenced done items, and current verification for artifacts
with mutation attempts. An empty ledger, a scaffold without verification, or a later write
that invalidated a passing check cannot verify closure. Read-only discovery does not require
artifact verification. This verdict does not certify that the ledger covers every requirement
of the human's goal; callers still need task-specific acceptance criteria.
A caller can bind a known candidate to one immutable delivery when continuing a session:
$input = [ 'prompt' => 'Finish the focus screen', 'session' => $sessionId, 'delivery' => json_encode([ 'workspace' => $candidateWorkspace, 'artifactPath' => 'src/Plugins/Owned/Services/FocusCounterView.php', 'test' => ['path' => 'tests/Plugins/Owned', 'filter' => ''], 'screen' => ['name' => 'focus', 'type' => 'focus-counter'], ], JSON_THROW_ON_ERROR), ];
CLI and HTTP accept delivery as a JSON string (--delivery on the CLI).
The DeliveryScope::parse() SDK also accepts a PHP array for direct use. The invocation records
session.delivery_declared with its observed caller provenance. Omitting delivery on later
turns retains it; an identical declaration is idempotent, and a different or malformed one is
refused before the turn runs. Use a new session for a different delivery. This input belongs to
the caller of agent, which is outside the resident's tool catalogue.
At each proven current final_answer with no pending question, the runtime reads AcceptanceEvidence again from the native stream,
current files and configured draft store. The result covers
scope: declared_delivery_and_recorded_work: current positive evidence can satisfy only the
identified producer's artifact, while open todos, unevidenced dones, explicit red judges and
other artifacts still count. A write after the delivery's test receipt blocks its closure,
even if a separate ledger verifier subsequently says green. No test is run during closure.
The returned closure and its session.closure_derived event contain the same sampled
observation and delivery declaration reference. These are historical observations, not locks,
approval, permissions or browser verification; a later turn must observe again. Sessions without
a declaration retain the original recorded-work verdict. All other termination causes, including unknown, produce no closure.
Evidence: greenhouse decisions/0390 and evidence/0708.
If the provider reports a truncated response, agent returns ok: false, truncated: true,
provider, outputLimit, stopReason, and the session id when one exists. It records no final answer or
closure for that incomplete response. Earlier effects and recorded usage remain in the session.
Containing what an agent may reach
An agent runs contained from the CLI, not only when a parent delegates to it. The withdrawal is a fact of the session, recorded in its stream — not a sentence in the prompt asking nicely:
# by name, when you know exactly which tools to take away php coa agent "review this app and report" --session=review --deny=plugins:enable,make # by effect class, which covers what a list of names forgets php coa agent "review this app and report" --session=review --denyEffects=mutating
Classes are mutating, external, irreversible and authority, resolved against the live
catalogue — an operation added tomorrow is covered the day it exists. An operation that never declared
its effects is denied, not waved through: unknown ranks above known-bad, so a catalogue nobody
classified withdraws entirely, and when that happens the command refuses and says so rather than
handing back a mute agent.
--deny needs --session: the option table lives in the session, and a prohibition that cannot be
recorded would not survive the first step.
Why a class and not a list: a measurement (settlement-q-p20p.md) put an agent under a task it could
not finish without mutating, took five tools away by name, and watched it reach for a sixth that
mutates — three times out of three. The list is worth exactly what whoever wrote it remembered.
The surfaces — where you drive it from
Console\Application is the single door of the CLI: coa on its own, a named command, the TUI, a
one-shot chat. Tui\AgentScreen renders the agent screen as text — the actor markers travel inside
the text, so a painter can colour by origin and the same screen still works where there is no colour.
Web\BoardPage renders the session's work as a live Kanban board in a browser: four columns,
and exactly one write — answering the question that paused the session, through two buttons
born disabled. They arm only when a token with the agent:answer scope is pasted; the token
travels in the Authorization header — never in a URL, never in browser storage — and the server
refuses any caller without a verified actor, showing the refusal verbatim. The page never folds
the stream client-side — the fold is agent:board, shared with the CLI — and when the live bridge
pushes a fact the page repaints the activity line and fetches the fold again, so reconnecting is
catching up. A card born already done is set apart, never animated as if it had crossed; a card
held by an open question sits in blocked saying why. Serve agent:board and agent:answer over
HTTP (config/http.php), point the page at your Mercure hub, and with no hub it says so instead
of pretending to be live.
Steering a session from any of them — agent:goal, agent:mode, skill:invoke
A session carries a standing goal — the human's intent, seeded from the first prompt and
changeable mid-session: agent:goal sets it, clears it, or reads it, over cli, tui, mcp and
http. The gate judges targets against it, and in auto mode it bounds what runs without asking;
the system prompt of every run speaks for the goal and the mode as they stand when that run starts.
agent:mode reaches the session over http too, so a Desktop's mode chip changes the real session,
not a label. A human runs a user-invocable skill with skill:invoke, which returns the skill's
body to put in front of the agent — including a skill marked disable-model-invocation, which the
model's own door, skill:load, refuses. All three are deliberately off the model's tool table
(AgentTable): a session must not widen its own standing ask, raise its own autonomy, or hand itself
a skill the human kept. And none of them pre-consents anything: a call that requires a signature, or
reaches a third party, still stops in every mode, whatever the goal names.
Growing the app — capabilities, capabilities:refresh, capabilities:enable
The capability→package index is derived from what the registry publishes, never written by
hand: every announcing package declares "type": "milpa-capability" on Packagist with its full
contract (extra.milpa.capability), and capabilities:refresh turns that into a dated artifact
under var/. Three authorities answer «what exists» and the rank is executed, not implied:
installed.json (what IS) over the derived index (what EXISTS, dated) over a small offline floor —
and every answer names which one it used. After capabilities:enable installs, what the registry
promised is compared with what arrived, and any difference is recorded: a package's
declaration about itself is a claim, not a classification.
Most of these exist because a measurement said they were needed, not because they seemed like a good
idea. The settlements live in the monorepo (docs/library/settlement-q-*.md) and the docblocks cite
which one.
Install
composer require milpa/app-runtime
A host composes it: this package boots nothing on its own and knows nothing about your app. It receives the session store, the operation catalogue and the model credential from whoever builds it — which is whoever holds the kernel.
Optional packages widen what it offers, and their absence is handled rather than assumed:
milpa/auth for token verification, milpa/data for persisting them, milpa/devtools for coa doctor, coa repair and coa update. Without them those surfaces are simply not offered — the app
never promises what it cannot do.
Passkey gate
One session, one scope, one middleware the panel names. PasskeyPlugin owns the whole passkey
ceremony — registration, sign-in, the session it mints — and registers PasskeyGateMiddleware in the
container under its own class name. A panel (milpa/admin, or any route of yours) puts identity in
front of itself by naming that class in its middleware list; it learns nothing about milpa/auth.
Identity lives where the ceremony lives (greenhouse decisions/0206).
The gate reads the session cookie the sign-in ceremony set and looks it up in the session store — the cookie value is never trusted on its own. Then:
| the request carries | a browser (GET accepting text/html) gets |
anything else gets |
|---|---|---|
| no live session (no cookie, unknown, expired, revoked) | 302 to /webauthn/signin?next=<where it was going> |
401 {ok:false, error:"unauthenticated", signin:"/webauthn/signin"} |
| a session without the scope | 403, a page: Authenticated, but the scope milpa.admin is not granted, naming the principal, with a Use another passkey link |
403 {ok:false, error:"scope_denied", scope} |
| a session with the scope | the route, with the AuthContext attached under milpa.auth (AuthenticateMiddleware::ATTRIBUTE) — signed in as passkey:<credential id> |
the same |
next is validated server side as a local absolute path: //evil, https://x and \x all fall
back to /. The sign-in page never redirects to a URL somebody else chose.
The operator sequence — from a fresh app to a panel that opens only for your key (run once with a physical YubiKey, greenhouse evidence/0519):
- Install what the door is made of. A fresh
composer create-project milpa/frameworkapp does not shipmilpa/auth— the ceremony, the session store and the gate middleware live there:composer require milpa/auth
Since 0.118 aPasskeyPlugindeclared without it refuses to boot and names this command; before, it mounted nothing and a panel naming the gate answered a mute500. Theidentity:*operations you need below are offered by the runtime on their own;milpa/agent+milpa/ai-gatewayare only for the agent operations. - Declare the plugin and the relying party. In
config/plugins.phplistMilpa\AppRuntime\Web\PasskeyPlugin::class; inconfig/app.phpdeclare'passkey' => ['rpId' => 'localhost']. TherpIdmust be the host the browser is on, and WebAuthn needs a secure context (https://, orlocalhost). WithoutrpIdthe plugin mounts nothing — a relying party nobody chose is one nobody can trust. With it, and no session store registered by the host, the plugin provides one (var/passkey/sessions.json). - Register the key. Open
GET /webauthn/enroll, press Register with passkey, touch the key. The page prints the credential id (base64url). The credential is now registered — the house holds its public key — but recognized by nobody: registering grants nothing. - Root the credential id out of band.
config/identity.php:<?php return ['rooted' => ['<credential id>']];
The root is read, never written, by the running app: the only way in is this file. - Enroll it with the panel's scope — a governed, signed operation:
php coa identity:enroll --fingerprint=<credential id> --scopes=milpa.admin --sign
How this is authorised today:identity:enrollis declaredrequiresConfirmation: true, so the CLI refuses it without--sign.--signsigns this exact call — operation, arguments, host — with your gpg key (the YubiKey through gpg-agent); the runner verifies the signature and hands the handler aGrantedAuthorization. The handler then checks that the grant coversidentity:enrollfor this fingerprint, that the id is inconfig/identity.php'srooted, and only then writes the recognition tostorage/identity/enrollments.jsonwithauthorized_by: key:<your fingerprint>.--scopesis an array argument — repeat the flag for more than one (--scopes=milpa.admin --scopes=agent:read). Overhttp/mcpthe operation additionally requires a caller holding theidentity:enrollscope. On the CLI a currently recognized signer's scopes are checked after signature verification;identity:enrollrequires that scope. A key never recognized retains the local bootstrap behavior. To make your gpg key the house's recognized root as well — so the ledger names it (authorized_by: bootstrap) and the same enrollment can run overhttp/mcp— bootstrap once, on an empty house, before rooting the credential:<?php return ['bootstrap' => true, 'rooted' => []]; // config/identity.php, first run only
php coa identity:bootstrap --scopes=identity:enroll --sign # one touch: your key becomes the rootidentity:bootstraprefuses oncerootedis non-empty or anything was ever recognized — it is a one-time act. Then write the credential id intorootedand enroll as above. Revoking (php coa identity:revoke --fingerprint=<credential id> --sign) laysrevoked_byover the entry and the sign-in list stops offering the key; enrolling the same id again re-admits it and keeps the revocation in the entry'shistory— the ledger records facts, it erases none (greenhouse decisions/0207). Active passkey sessions use the enrollment's current scopes on every request: reducing permissions takes effect immediately without requiring a new sign-in. An empty scope list preserves authentication and grants no scoped access; revocation ends the session. - Name the gate. Where the panel's middleware is declared (
admin.middlewareformilpa/admin, themiddlewareof anyRouteof yours):'admin' => ['middleware' => [Milpa\AppRuntime\Web\PasskeyGateMiddleware::class]],
- Sign in.
GET /milpa/admin→302to/webauthn/signin?next=/milpa/admin→ Continue with a passkey → touch → the cookie is set and the browser returns to the panel,200.
If Continue with a passkey does nothing — no dialog, no error, the button stays disabled — a browser
extension has most likely replaced navigator.credentials.get (password managers that offer their own
passkeys do; the console shows the extension's content script). The pages now say so before waiting on
the call; retry in a browser profile without that extension (greenhouse evidence/0519).
Why the sign-in page works with a hardware key: enrollment registers a non-discoverable credential
(residentKey: discouraged, so a key with scarce slots is not consumed), and a browser only finds one of
those when the request names it. The authentication and intent options therefore return
allowCredentials with every credential id that is registered AND enrolled — POST /webauthn/register
stays open and registering grants nothing, so a key nobody enrolled is never offered; an id is not a
secret, the private key is. The intent page (the D-01 approve ceremony) now requests
userVerification: 'required', the same bar the enrollment ceremony sets (greenhouse evidence/0486).
Config keys (config/app.php, under passkey):
| key | default | what it decides |
|---|---|---|
passkey.rpId |
none — required | the relying-party id every assertion binds to; without it, no routes |
passkey.cookie |
milpa_session |
the cookie the session id travels in (HttpOnly, SameSite=Strict) |
passkey.ttl |
3600 |
session lifetime in seconds, from the moment the ceremony mints it |
passkey.sessions |
<root>/var/passkey/sessions.json |
where the provided FileSessionStore writes — ignored when the host registered its own SessionStore |
passkey.gate.scope |
milpa.admin |
the one scope PasskeyGateMiddleware requires (the * wildcard an identity:bootstrap root holds also opens it) |
POST /webauthn/register stays open: registering grants nothing, enrolling is the act, and the root gate
is the file only you write.
The session on the operations surface
The same session is a principal of the operations surface (greenhouse decisions/0208). Once the door
is wired, PasskeyPlugin also registers Milpa\AppRuntime\Web\PasskeySessionMiddleware — under its own
class name, and as the container's Milpa\Auth\Contracts\AuthContextFactory (the same instance; a host
that registered its own factory keeps it). It reads the cookie and puts the resulting AuthContext under
milpa.auth, exactly where AuthOperationHttpPolicy judges every operation that declares scopes — so
a browser signed in with a key holding agent:run can POST /agent, and one without it gets 403.
- Precedence — the Bearer decides. If
AuthenticateMiddlewarealready left a context that is authenticated or invalid, the request passes through untouched: a rejected Bearer is never laundered by a cookie. Only an absent or anonymous context lets the cookie speak. - Revocation parity. The cookie is worth what it is worth at the panel's door: resolved through
milpa/auth'sStartSessionand re-checked against the enrollment ledger on every request. A revoked passkey's session is destroyed, the response carries an expiringSet-Cookie, and the request goes on anonymous — no error from the middleware; the operation's policy answers401if it needed an actor. The ledger judges onlypasskey:*principals: a session the host minted itself throughmilpa/authunder the same cookie (token:…,user:…) is attached as it is and never destroyed here. - CSRF posture. A mutating request (
POST,PUT,PATCH,DELETE) authenticates from the cookie only when itsContent-Typeisapplication/json(parameters allowed) and, ifSec-Fetch-Siteis present, it sayssame-originornone. Otherwise the cookie is ignored — not even read — and the request continues anonymous.GET/HEAD/OPTIONSauthenticate from the cookie unconditionally.
Compose it in public/index.php after the Bearer middleware and before the handler (milpa/framework
ships this composition as App\Http\IdentityChain, executed by its own test):
use Milpa\AppRuntime\Web\PasskeySessionMiddleware; use Milpa\Auth\Contracts\CredentialVerifier; use Milpa\Auth\Http\AuthenticateMiddleware; use Psr\Http\Message\ResponseInterface; use Psr\Http\Message\ServerRequestInterface; use Psr\Http\Server\MiddlewareInterface; use Psr\Http\Server\RequestHandlerInterface; $container = $kernel->container(); $chain = []; if ($container->has(CredentialVerifier::class)) { $chain[] = new AuthenticateMiddleware($container->get(CredentialVerifier::class)); // the Bearer decides first } if ($container->has(PasskeySessionMiddleware::class)) { $chain[] = $container->get(PasskeySessionMiddleware::class); // the cookie speaks when it said nothing } foreach (array_reverse($chain) as $middleware) { // nest: first declared runs first $handler = new class ($middleware, $handler) implements RequestHandlerInterface { public function __construct(private readonly MiddlewareInterface $m, private readonly RequestHandlerInterface $next) {} public function handle(ServerRequestInterface $r): ResponseInterface { return $this->m->process($r, $this->next); } }; } $response = $handler->handle($request);
Upgrading
events:catalogue answers for the APP, not for the process
An emitter declares its events to the dispatcher when it is constructed, so a CLI process that builds
almost none of them heard about almost none of them: measured on fresh cattle, the catalogue listed 7 of
the family's 24 framework events (greenhouse decisions/0228, second slice). The package now speaks for the
emitter nobody built — it names a Milpa\Interfaces\Event\DeclaresEvents holder in its own manifest, and
this fold reads those manifests and declares on the emitter's behalf, so events:catalogue and
house:context's events section both answer for what the app HAS installed.
{ "extra": { "milpa": { "events": ["Acme\\Shop\\Event\\ShopEvents"] } } }
- The floor moved:
milpa/core >= 0.12(whereDeclaresEventslives), and with itmilpa/runtime >= 0.14,milpa/plugin >= 0.17andmilpa/live >= 0.23— the versions whose manifests name their holders.milpa/mcp-server >= 0.7andmilpa/admin >= 0.16do the same for the apps that install them. - The manifests read are
vendor/composer/installed.json— what Composer really resolved — plus the app's owncomposer.json, so an app declares the events IT dispatches the same way. Nothing is probed and no path is invented; an app without aninstalled.jsonanswers from its dispatcher alone. - The dispatcher stays the authority. It keeps the first declaration of a name, so an emitter that was really constructed is never overridden by its manifest, and asking twice changes nothing.
- A new
warningslist: a manifest entry naming a class that is not autoloadable here, or one that is not aDeclaresEventsholder, comes back as{package, class, why}withokstilltrue. Its events are MISSING from the catalogue, which is exactly why it is said out loud instead of dropped — and no row is invented for it.
events:catalogue — the house counts its own events; house:context gains a section
events:catalogue answers what this app's dispatcher was told exists against what it really
dispatched in this process (greenhouse decisions/0228), and house:context carries the same fold
compact under a new events key. Both read the dispatcher and nothing else — the manifest pass described
above came after this and declares TO that same dispatcher: the authority on «what events exist» is the
emitter, and the one place every dispatch passes through is the dispatcher.
- The floor moved:
milpa/core >= 0.11, which is where the contract lives (Milpa\Interfaces\Event\DeclaredEvents,EventDeclaration).MilpaEventDispatcherInterfacewas not widened — a dispatcher either implements the new interface or is asked nothing. - An app on an older dispatcher is not broken, it is named.
milpa/events < 0.4implements no memory of what was declared or dispatched, so both answerok:falsewith the dispatcher's class and the interface it lacks — never an empty list, which would read as «this app dispatches no events».composer update milpa/eventsto>= 0.4and the rows appear. - A name dispatched without a declaration is listed as debt, with
declared: falseandnullfor everything only a declaration could say. Nothing is invented to fill the row out, and nothing is hidden for lacking one.
Scoped plugin authoring
The host registers one PluginAuthoringPolicy as a tool CallPolicy and an OperationBoundary.
CLI, MCP and HTTP executions carry their current ToolContext into the runner; agent, sequence and
recipe drivers preserve it when opening the next door. InvocationContext remains attribution,
not permission. These contracts require milpa/tool-runtime >= 0.17 and milpa/console >= 0.20.
A finite caller needs the exact scope plugins.<Plugin>:write to author that plugin. For example,
plugins.Owned:write permits writing src/Plugins/Owned/ and tests/Plugins/Owned/. This follows
Permission's namespace/resource/action spelling; it does not expand roles, accept globs or assign
meaning to the experimental plugin:Owned string. Activation still needs its own authorization.
make,implementandeditrequire a canonical plugin name. Implementations must currently use either one complete body orimplement'smode=start,append, andfinishprotocol. Each section still runs in a confined trial and must be promoted before the next call can use it. Parts remain beside the scaffold as.php.milpa-part, inside the same plugin write set; they never replace executable PHP untilfinishpasses the existing verification gate and its trial is promoted. Revocation also blocks promotion of pending parts.testrequires a relative path undertests/Plugins/<Plugin>/. Tests and verifier subprocesses run in the same write boundary, with read-only root/vendor, private trial state and temporary storage, an ephemeral PHPUnit cache, and unshared network/PID namespaces. Missing confinement refuses execution; it never falls back to writing the host.- Trial stdout and stderr are drained together, so a verbose warning cannot block the child behind an unread pipe. Both channels and the exit status are retained, including output produced before the existing trial deadline kills an unfinished process.
- A failed native
teststays unsuccessful. Its error text is JSON with schemamilpa.trial-test-failure/v1,ok: false,ran_in_trial: true,applied: false, the workspace,trial_exit, the original structuredoutput(ornull), and separatestderr. This survives the tool channel's exception and the durable session record. Missing output does not imply a PHPUnit verdict; unknown producer counts remain unknown. Invalid UTF-8 in diagnostics becomes the Unicode replacement character. DirectToolResult.dataconsumers keep the original producer data. No promotion instruction is added to failed tests. - A successful trial is a proposal.
sandbox:promoteandsandbox:undojudge every affected file against the authority of the current call before writing the first one. Mixed resource exports, traversal and symbolic links refuse as a whole. A saved trial never saves permission to export. - The diff compares the trial copy with its original host manifest. Changes made only in the host are not trial edits; stale checks separately reject conflicts on files the trial actually changed. This lets a current regrant authorize a pending trial without overwriting the enrollment ledger.
- Other mutating operations with empty declared scopes refuse finite callers. The agent's session
bookkeeping declares
agent:run. The local*mode keeps its existing behavior.
This confines authoring writes, including code run by a verifier. It does not isolate a local shell owner, hide readable files, make already activated host plugins untrusted, or provide transactional isolation against external writers. Authoring a plugin and activating its code are separate steps.
0.120.0 — driving the agent requires agent:run; the passkey session is a principal
The four operations that drive the agent — agent, skill:invoke, agent:goal, agent:mode — now declare
scopes: ['agent:run'] (greenhouse decisions/0208). Over HTTP the policy is consulted where before it was
not: an anonymous POST /agent now answers 401, and an authenticated actor without the scope 403.
The CLI now checks declared operation scopes when MILPA_TOKEN presents a verified identity with
nonempty scopes. It uses the same PolicyGate::authorizeScopes judgement as the agent door, before
asking for a signature or session consent. Consent cannot supply a missing scope; a sufficient scope
does not replace consent. This requires milpa/tool-runtime >= 0.16.
With --sign, the current verified GPG signer also carries its recognized scopes into the shared
gate and delegated tools. The enrollment ledger takes precedence over static policy: revoked entries
have empty authority, and an unreadable ledger refuses. A key never recognized retains the existing
local bootstrap behavior. An explicit signature authenticates read operations as well; a stored
session owner never supplies that authority. When a token and signature are both presented, their
scopes intersect, so signing cannot widen the token.
HTTP agent turns additionally require milpa/console >= 0.19: the authenticated request's tool
authority travels separately from InvocationContext, through the runner to the agent's governed
door. A passkey or Bearer caller keeps its own scopes, including an empty list; the server's
MILPA_TOKEN cannot replace them. A web turn missing that authority refuses before calling a model.
Absent, invalid and empty-scope tokens retain the local process’s * default (greenhouse decision
0311); finite callers are additionally subject to the plugin authoring policy below. MCP over stdio
also retains its * context. An MCP client that authenticates as a principal of its own needs agent:run
for skill:invoke, agent:goal and agent:mode, as it already did for agent:sessions and agent:show.
The * wildcard an identity:bootstrap root holds keeps admitting. An app that exposes any of the four
in config/http.php — by name, or through expose: ['*'], which includes them — without an
OperationHttpPolicy now refuses to boot, as it does for every scoped operation.
- Tokens: mint them with the scope —
php coa token:new desktop --scopes=agent:run(addagent:read/agent:answerfor the session reads and the gate answers, as before). - Passkeys: enroll them with it —
php coa identity:enroll --fingerprint=<credential id> --scopes=agent:run --sign(repeat--scopesformilpa.adminif the same key opens the panel). - The session itself:
PasskeyPluginnow registersPasskeySessionMiddleware(also as the container'sAuthContextFactory), and it only counts oncepublic/index.phpcomposes it afterAuthenticateMiddleware— existing apps copy the new composition frommilpa/framework'spublic/index.php(see The session on the operations surface above). Without that line the passkey cookie keeps opening the panel only. A house withoutmilpa/dataalso needs itsOperationHttpPolicyregistered withmilpa/authalone (the skeleton'sconfig/boot.phpnow does, throughApp\Http\IdentityWiring): a policy gated on the token store leaves the cookie-only house with no policy to consult, and a scoped operation exposed inconfig/http.phprefuses to boot.
0.119.0 — the identity ledger keeps history
storage/identity/enrollments.json is a ledger of facts, not of state (greenhouse decisions/0207).
Enrolling a key that already has an entry — revoked or live — no longer overwrites it: the state it
replaces is pushed onto the entry's history list (most recent last) and the new {scopes, authorized_by}
becomes the live state, so a revocation is never erased by the recognition that follows it. Re-enrolling a
revoked key is allowed, under the same signed, rooted authority as enrolling. identity:enroll now says
what it did: history_entries (prior states kept for the key; 0 on a first enrollment) and, when the
standing entry was revoked, previously_revoked_by. Reads are tolerant: a ledger written before reads
identically, scopesFor (live state; revoked → null) and isEmpty (sealed by any entry) keep their
contracts, and history appears the first time a key is re-written. No migration. And a write the store
cannot make — the file cannot be opened, the disk refused the bytes, or the ledger holds content the store
cannot read — is refused rather than reported on: identity:enroll, identity:revoke and
identity:bootstrap answer ok: false naming the cause, nothing is written over unreadable content, and
such content is not a greenfield for identity:bootstrap.
0.118.0 — PasskeyPlugin refuses to boot without milpa/auth
A PasskeyPlugin listed in config/plugins.php while milpa/auth is not installed used to boot
quietly and mount nothing; a panel naming PasskeyGateMiddleware then answered a 500 that blamed
nothing (greenhouse evidence/0519). Now boot() throws a RuntimeException that names the fix:
composer require milpa/auth # or remove the plugin from config/plugins.php
No behaviour changes for a house that has the package. The sign-in, enrollment and intent pages also say when a browser extension replaced the WebAuthn API on the page, before waiting on it.
0.45.0 — capabilities:enable --dry-run requires a signature
--dry-run used to run without consent. It no longer does: the rehearsal now carries the full
operation's ceiling, so it asks like any other governed effect.
# before coa capabilities:enable milpa/devtools --dry-run # now coa capabilities:enable milpa/devtools --dry-run --sign
The exemption came from a descent — a declaration that the rehearsal reaches no further than the disk — and it was switched off because nothing could check it. The claim rests on the network, and the network here is observed by difference, which cannot tell does not reach out from reaches out and swallows the error. A ceiling lowered by a promise nobody can verify is worse than the nuisance of asking. The descent returns when it can be certified.
--dry-run still does exactly what it did; only the exemption is gone.
License
Apache-2.0 · © Rodrigo Vicente — TeamX Agency
Milpa is designed, built, and maintained by Rodrigo Vicente - TeamX Agency.
Review a screen before activation
With LivePlugin enabled and live.secret configured, /live/review lets an
identified reviewer create immutable screen revisions, compare their baseline and
proposal, try the proposed screen, activate the exact revision, and restore its
baseline. Reload an active page to load its new declaration. Existing open tabs
keep their signed view until reloaded; activation does not replace application code.
The operations screen:draft (name, type, props), screen:review (optional
revision), screen:promote and screen:rollback (required revision) use the same
revision service. They declare milpa:component:screen-review:draft, :read,
:promote and :rollback respectively. The review page requires :read; its
buttons enforce the corresponding action scopes. The component wildcard :*
grants all four in the component UI; assign the explicit scopes above to operation
callers. A shareable review URL is /live/review?revision=<id> under the
configured live route. Identity and scopes are still required.
The resident agent creates revisions in the host's review store. These three
revision mutations use their own lifecycle instead of the generic file trial;
scope and consent checks still apply. A scoped launch grant such as
--grant=screen_draft:name=todos can consent to proposals for that screen without
granting activation. Reading a generated revision does not require its hash to
have appeared in the original request. Plugin authoring and screen:declare
retain their existing trial routing (Greenhouse 0330).
Apps explicitly opt component types into ScreenPreviewRegistry during plugin
boot. A factory receives a PreviewEnvironment containing the immutable revision
ID, an isolated codec, and empty component and renderer registries. It must build
its complete component graph and HTML renderers with that codec and separate
persistence/effect collaborators. No active registry or renderer is used as a
fallback. See Greenhouse's complete ToDo example for an implementation with a
separate SQLite file per revision and private records per principal.
Preview is a trusted application factory boundary, not an operating-system sandbox for arbitrary PHP or external services. Preview actions still require the component's normal scopes. Test records are never copied to active storage. Revisions and their test stores remain under the app's ownership; this release does not delete them according to a retention policy.
Revisions live under var/screen-drafts; each file is addressed and verified by its
canonical content hash. Activation compares the reviewed baseline under a file lock,
then atomically replaces only that screen's declaration. The app's src, config,
public and composer.lock fingerprint must still match. Dependency selection is
bound through the lock; edits made directly inside vendor are outside this guard.
Changing app PHP, configuration, assets or dependencies requires creating a new
revision. Malformed active declaration stores are refused instead of overwritten
as empty stores. Promotion and restoration change screen declarations only, never
PHP, migrations, or application records. Greenhouse decision 0329 defines this cut.
Observed progress
Native trial calls record a host observation of the file changes made by that execution, separate
from the operation's permission ceiling. The session links that observation to its tool call;
repeated content in another trial does not create new progress. Promotion has a separate application
identity. A successful native test contributes behavioral evidence for its selector and input tree,
Changing only the timeout does not create another proof.
A known empty observation does not reset recovery. An unavailable observation neither clears a
pending recovery nor proves its window exhausted. Older producers without observations retain the
legacy session interpretation. This protocol requires milpa/agent >=0.44 when the optional agent
capability is installed (greenhouse decisions/0346, evidence/0663).
Trial input freshness
A repeated call reuses its trial plan only while the copied host inputs still match the original
manifest. Source additions, edits, deletions and undo renew the plan; top-level var/, live-mounted
vendor/ and .env are outside that copied-input comparison. TrialWorkspace::hasCurrentInputs()
checks all copied inputs, while stale() continues to check only a proposal's promotion targets.
Renewal retains the old workspace and pending diff under the existing 24-trial retention bound; it never rebases or promotes that proposal. Unreadable baselines cannot establish freshness. This does not provide an atomic snapshot against concurrent host writers or change the independent repeated-failure guard. Greenhouse decisions/0347 and evidence/0664 measure the native path.
Read a candidate before continuing
candidate:state reads a single-file edit or implement candidate from the session's native
receipts and current files. It is offered by AgentOperations when the agent capability is
installed; no model connection or extra app provider is needed. The operation requires
agent:read and declares read-only effects.
php bin/coa candidate:state --session=repair-session --workspace=w1234567890abcdef --json
The same projection is available to PHP consumers:
use Milpa\AppRuntime\Agent\CandidateState; $state = CandidateState::read($appRoot, $sessionStore->stream($sessionId), $workspaceId);
Pass an absolute, resolved application root and the trusted session stream. The projection returns
pending, promoted, contradicted or indeterminate, with a reason, the candidate path/hash
when known, and references to the producing events and physical receipts. It reads again on each
call; it has no durable state or cache of its own.
A pending candidate has a matching trial copy and current copied-input baseline. A promoted
candidate has an exact promotion receipt, matching host bytes and a collapsed copy. A contradiction
or insufficient evidence offers no continuation. verification.scope=producer_declaration
preserves what the producer reported: syntax/conformance alone does not become behavior acceptance.
authorization is always not_evaluated. When present, next describes the existing
sandbox:promote operation and its workspace, with requiresRecheck=true. Re-read immediately
before proposing it through the normal governed runner. The read neither grants permission nor
reserves future bytes; the normal native gates still decide execution. Reading through the agent
continues to record ordinary tool-call events.
This contract covers one added or modified file. Pending reads use the native copied-input domain; promoted reads cover the files recorded in the baseline, not later added inputs, vendor or the process environment. It does not certify all execution dependencies, test/review acceptance or readiness to deploy. Evidence: greenhouse decisions/0375–0376 and evidence/0692–0693.
Join candidate, test and screen evidence
acceptance:evidence is an agent:read operation offered with the agent capability, without a
model connection. It joins the candidate's native receipts and current files, the latest test
attempt in that session, and its latest immutable screen review. The caller supplies the question:
use Milpa\AppRuntime\Agent\AcceptanceEvidence; use Milpa\AppRuntime\Web\ScreenDrafts; $evidence = AcceptanceEvidence::read( $appRoot, $sessionStore->stream($sessionId), $candidateWorkspace, ['path' => 'tests/Plugins/Owned', 'filter' => ''], ['name' => 'focus', 'type' => 'focus-counter'], $container->get(ScreenDrafts::class), );
Pass the host's configured ScreenDrafts service; pass null when unavailable. The native
operation uses that same service and SDK. Its arguments are session, workspace, test
(the exact path/filter object) and screen (name/type, optionally an exact definition).
The workspace is the edit/implement candidate, not the later test workspace.
Catalogue queries (screen_review with {} or {"revision":""}) do not replace the latest
directed review. The collector and receipt join select by request intent, never success: a
later failed or malformed directed attempt stays relevant and cannot recover an earlier green
result. A catalogue alone supplies no exact review. Current files and the selected revision
are still re-observed; catalogue queries do not freeze evidence freshness.
The versioned milpa.acceptance-evidence/v1 result distinguishes current_evidence,
historical_evidence, failed, incomplete and indeterminate. Test outcome, ran, known
counts, actual scope and coversRequestedScope remain separate: a failed or filtered attempt
does not stand for the entire requested suite. Unknown counts stay null, and zero stays zero.
Changing copied inputs invalidates failed evidence as well as passing evidence. The reader
uses complete durable receipts, not a model-window preview, and re-observes current files.
Full output and stderr remain in the session stream. test.receipt identifies the original
session.tool_called result by toolCallSeq, character count and SHA-256 of its UTF-8 bytes.
This compact projection preserves the verdict without repeating the diagnostic; it is not a new
receipt store or a guarantee that arbitrary user-supplied screen definitions fit every model window.
A refusal without a trial remains indeterminate and does not acquire a PHPUnit verdict.
authorization=not_evaluated and humanApproval=not_recorded are unconditional. A screen review
is a read, not human acceptance. This does not activate a screen, deploy, certify browser behavior,
lock future bytes, or certify execution inputs outside the native copied files and screen build.
Concurrent changes can invalidate any later action. Output bytes that cannot be reconstructed
exactly from the native runner's JSON record plus LF remain indeterminate; extra stdout is not
silently normalized away. Evidence: greenhouse decisions/0381 and evidence/0698.
Run termination and closure
The agent result includes termination: {reason, receipt} for attempts that reach the model
invocation. The same observation is appended as session.run_terminated when a session event
store is available. Reasons come from milpa/ai-gateway 0.25.0's base loop; unknown means the
current invocation has no proven producer observation. Answer text never supplies the cause.
Closure is derived only for a current final_answer with no pending session question. Refusal,
confirmation, blocking, exhaustion, stalled progress, declared house debt, invalid response and
exceptional exits cannot derive a closure. A genuine final answer may still have paused: true
if the host recorded a question; it has no closure. A final cause alone certifies no completed
work, permission or approval: the existing recorded-work verdict remains the judge.
The protected ask() and orchestrator() signatures remain unchanged. Overriding the factory
while preserving the base ask(), orchestrator run() and termination() methods retains
provenance. Replacing any of those three methods yields unknown, even if a previous base run
left an observation. A reused producer must emit a new observation in the current invocation.
Older stale vendors without the API also yield unknown; Composer requires gateway >=0.25
when that optional capability is installed. Early refusals before ask() emit no run observation.