milpa / ai-gateway
Dual-provider LLM gateway for the Milpa PHP framework: OpenAI + Anthropic chat-completions client with tool-call format translation, an agentic tool-use loop, and a ToolRegistry facade for LLM/MCP exposure.
Requires
- php: >=8.3
- guzzlehttp/guzzle: ^7.10
- guzzlehttp/psr7: ^2.7
- milpa/core: >=0.6.2 <1.0
- milpa/tool-runtime: >=0.17.1 <1.0
- psr/http-client: ^1.0
- psr/http-factory: ^1.1
- psr/http-message: ^2.0
- psr/log: ^3
Requires (Dev)
- friendsofphp/php-cs-fixer: ^3.65
- phpstan/phpdoc-parser: ^2.3
- phpstan/phpstan: ^2.1
- phpunit/phpunit: ^11.5
Suggests
None
Provides
None
Conflicts
None
Replaces
None
- dev-main
- v0.25.0
- v0.24.5
- v0.24.4
- v0.24.3
- v0.24.2
- v0.24.1
- v0.24.0
- v0.23.0
- v0.22.2
- v0.22.1
- v0.22.0
- v0.21.0
- v0.20.0
- v0.19.0
- v0.18.1
- v0.18.0
- v0.17.2
- v0.17.1
- v0.17.0
- v0.16.1
- v0.16.0
- v0.15.0
- v0.14.1
- v0.14.0
- v0.13.0
- v0.12.0
- v0.11.0
- v0.10.0
- v0.9.1
- v0.9.0
- v0.8.3
- v0.8.2
- v0.8.1
- v0.8.0
- v0.7.0
- v0.6.0
- v0.5.0
- v0.4.2
- v0.4.1
- v0.4.0
- v0.3.1
- v0.3.0
- v0.2.3
- v0.2.2
- v0.2.1
- v0.2.0
- v0.1.1
- v0.1.0
- dev-rodrigoteamx/typed-termination
- dev-rodrigoteamx/structured-window
- dev-rodrigoteamx/enforce-withdrawal
- dev-rodrigoteamx/bounded-recovery
- dev-rodrigoteamx/lazy-tool-contracts
- dev-rodrigoteamx/output-contract
- dev-feat/stream-usage-token-cost
- dev-fix/flake-retry-twice
- dev-feat/context-self-healing
- dev-fix/provider-flake-retry
- dev-fix/null-tool-args
- dev-feat/intra-leg-budget
- dev-fix/notice-role
- dev-feat/forced-choice
- dev-feat/gateway-resilience
- dev-feat/reasoning-observer
- dev-feat/lazy-toolbox
This package is auto-updated.
Last update: 2026-09-14 22:53:18 UTC
README
Milpa AI Gateway
A dual-provider LLM gateway for the Milpa PHP framework — one client for OpenAI and Anthropic chat completions, translating each provider's tool-call wire format to and from a single shape, plus an agentic tool-use loop that drives a
milpa/tool-runtimeToolRegistry(resolve → validate → authorize → execute → audit) until the model is done.
milpa/ai-gateway is the LLM tier of Milpa: the piece that turns a milpa/tool-runtime
ToolRegistry into something a model can actually drive. LlmService implements
milpa/core's LlmServiceInterface seam against two concrete providers — OpenAI and
Anthropic — so callers write one message shape and one tool-call shape regardless of which
provider answers. AgentOrchestrator runs the loop every agent needs: ask the model, execute
whatever tools it asks for through the registry pipeline, feed the results back, repeat until
the model returns a final answer or a step budget runs out. No product coupling, no
Telegram/HTTP-specific code — those live in your host application.
Install
With AgentOrchestrator(lazyTools: true), describe_tool lists discoverable names and
descriptions. Describing a tool makes it callable with its full input schema on subsequent
requests. Undiscovered tools are not advertised as callable empty-object signatures. Discovery
does not authorize execution: the same tool registry and gates still judge every real call.
An empty catalogue exposes no discovery tool. The default full catalogue is unchanged.
composer require milpa/ai-gateway
Semantic recovery
An optional ProgressProbe owns the evidence and the recovery window. Its additive recovery
field reports pending, recovered, or exhausted. Pending recovery survives successful tool
calls and missing observations; only measured recovery clears it. Exhaustion returns
AgentOrchestrator::PROGRESS_STALLED with the producer's receipt before another model call.
Preparation can therefore lead to real work without making repeated bookkeeping count as work.
The total step budget remains in force. Probes omitting the field keep their previous behavior;
confirmation, debt and abandonment retain their exits, and abandonment does not reset recovery.
Quick example
Register a tool on a ToolRegistry (from milpa/tool-runtime), wrap it in McpClientService,
and hand both to AgentOrchestrator along with an LlmService:
use Milpa\AiGateway\AgentOrchestrator; use Milpa\AiGateway\LlmService; use Milpa\AiGateway\McpClientService; use Milpa\ToolRuntime\ToolRegistry; use Psr\Log\NullLogger; $registry = new ToolRegistry(new NullLogger()); $registry->register( 'get_time', 'Get the current time', [], fn () => ['time' => '12:00 PM'], ); $mcpClient = new McpClientService($registry); $llm = new LlmService(apiKey: getenv('OPENAI_API_KEY'), model: 'gpt-4o', provider: 'openai'); // LlmService talks HTTP through PSR-18 (Psr\Http\Client\ClientInterface), defaulting to a // Guzzle client when none is injected — see "Bringing your own HTTP client" below. $orchestrator = new AgentOrchestrator($llm, $mcpClient); echo $orchestrator->run('What time is it?'); // -> asks the model, the model requests `get_time`, AgentOrchestrator executes it through // the registry, feeds the result back, and returns the model's final answer.
Swap provider: 'anthropic' and a Claude model name (or let a claude model name in
$model select it automatically) to point the same call at Anthropic instead — LlmService
translates the tool list, the message history, and the tool-call response to and from
Anthropic's shape internally, so AgentOrchestrator and McpClientService never see a
provider-specific format.
Run termination
run(): string keeps its existing response text and exceptions. After calling the base
loop, read $orchestrator->termination() for the producer's RunTermination observation:
use Milpa\AiGateway\RunEnd; try { $answer = $orchestrator->run('Inspect the app.'); } finally { $end = $orchestrator->termination(); // $end?->reason === RunEnd::FinalAnswer means the final-answer branch returned. // It does not mean the requested work is complete or approved. }
reason distinguishes final_answer, tool_refused, confirmation_required, blocked,
steps_exhausted, progress_stalled, house_debt, invalid_response, interrupted,
output_truncated, and failed. For progress_stalled, receipt carries the probe's original
receipt. toArray() exports both fields. A refusal and a final answer can have identical text;
the cause comes from the executed branch, never from matching that text.
The observation resets to null when each base run begins and is available after a return or
an exception escapes; the same exception object still propagates. Captured observations retain
their values when the instance is reused. The getter describes the latest base-loop run;
subclasses that replace run() must not present an earlier base-loop observation as a new one.
Concurrent or reentrant runs on the same orchestrator are not supported. Session questions,
completion evidence and authorization remain the host's responsibility.
The agent loop
AgentOrchestrator::run() alternates between two calls until the model is done or
$maxSteps (default 20) is reached:
- Ask —
LlmService::generateResponse()sends the running message history plus the registry's tool summaries (McpClientService::getToolSummaries()) to the provider and returns a single OpenAI-shaped assistant message. - Act — if that message carries
tool_calls, each one is executed viaMcpClientService::callTool(), which runs it through the fullToolRegistrypipeline (validate → authorize → execute → audit) under whateverToolContextwas set withsetToolContext(). The result — rendered through aRendererRegistrywhen one is configured, JSON otherwise — is fed back into the message history as atoolmessage, and the loop repeats.
If a tool result requires confirmation or is blocked by policy, the loop stops immediately and returns that outcome instead of continuing — the caller (a chat handler, a CLI, a bot) is responsible for the confirm/cancel round trip on the next user turn.
generateResponse(maxTokens: 4096) requires a positive output limit. It sends
max_completion_tokens to OpenAI-compatible endpoints and max_tokens to Anthropic.
Oversized structured tool exceptions use a bounded milpa.tool-failure-window/v1 projection.
Its outer ok:false describes the failed call; preview preserves the decoded error where it
fits, and partial plus omitted identify removed fields by JSON Pointer, character count and
SHA-256. original identifies the complete error by size and hash. The session recorder keeps
that original receipt; the projection does not write or replace it. Large object string fields
may be omitted, while arrays and scalar facts are never rewritten. If the structure still does
not fit, preview:null and a root omission report that it is unavailable. The same applies
when number spellings cannot survive PHP decoding and encoding unchanged. A preview is not a
complete receipt or authorization. Short errors, unsupported JSON and successful results keep
their existing paths; no context or per-result budget is increased.
When the provider reports output truncation (length / max_tokens), buffered and SSE
responses throw OutputTruncatedException before any tool call from that response can execute.
Anthropic's model_context_window_exceeded is also treated as truncation. The exception exposes
provider, stopReason and the requested maxTokens; it does not expose partial tool
arguments. Usage is still observed. There is no automatic retry, including during the guided
retry of a degenerate answer. Previously completed steps remain completed.
Stopping the loop before it acts: ToolCallGate
The loop above executes what the model asked for. ToolCallGate is the seam that lets somebody
else decide, before each call, whether it proceeds — and ToolCallRecorder is told after, with what
the tool answered. Both contracts live in milpa/tool-runtime (Milpa\ToolRuntime\Gate), where
every caller of tools already depends: a model is one caller, a governed door opened by a human is
another, and both ask the same question (greenhouse decisions/0225). This package keeps the two names
as interfaces that extend the base, so whoever implemented them here is still a gate or a recorder
wherever the base is asked for.
use Milpa\ToolRuntime\Gate\ToolCallGate; use Milpa\ToolRuntime\Gate\ToolCallRecorder; interface ToolCallGate { // The reason this call does not proceed, or null if it does. public function refuse(string $tool, array $arguments): ?string; } interface ToolCallRecorder { // Told after the call, with the rendered result and whether it succeeded. public function recorded(string $tool, array $arguments, string $result, bool $ok): void; }
refuse() runs before the call and recorded() after it. The two halves are not symmetric on
purpose: refusing is about INTENTION (what the caller is about to do), recording is about OUTCOME
(what happened). A refusal is not a tool error: McpClientService::callTool() throws
ToolCallRefusedException (which extends tool-runtime's ToolCallRefused) and the orchestrator
catches it apart, before any generic catch, and ends the turn. The gate is an interface here and
nothing else: this package brings no policy of its own — milpa/app-runtime implements it with the
session's permissions, the autonomy mode and the signatures.
The table: OptionTable
Since 0.5 the loop can operate against a table — the set of options the agent currently has in front of it — through one port with two questions that look alike and are not:
interface OptionTable { public function remove(string $option, string $code, ?string $message = null): void; public function removed(): array; // the PROJECTION: what the model sees public function wasRemoved(string $option): bool; // the FACT: did an authority declare it gone? }
The catalogue the model sees is re-derived every step from the registry minus removed() — it
is a projection, never a snapshot. And a refusal that removed the option no longer ends the run:
there is nothing left to route around, so the reason goes back to the model and the loop continues
with a different table. Both came out of measurement: a frozen catalogue made "did the agent re-read
the world?" unanswerable, and a run-ending refusal turned every denial into a shutdown (0 of 32
runs recovered).
SecondOpinionGate also stopped being quiet anywhere: an empty answer, a verdict-less answer and a
failed judge each say so with their own cause, a denial leaves a witness on stderr outside the
stream it writes to, and an ALLOW logs nothing — approving silently is correct; failing silently
was being counted as approval.
Provider translation
LlmService speaks one shape to its callers — OpenAI's messages / tool_calls — and
translates both directions for Anthropic:
- Outbound:
systemmessages become Anthropic's top-levelsystemparameter;toolrole messages becomeusermessages carrying atool_resultcontent block; an assistant message withtool_callsbecomestool_usecontent blocks. Tool summaries are reshaped from{name, description, inputSchema}to Anthropic's{name, description, input_schema}, with an emptypropertiesobject substituted where a tool declares none (Anthropic requires a non-empty schema object, not an empty array). - Inbound: Anthropic's
content: [{type: text, ...}, {type: tool_use, ...}]array is flattened back into a single OpenAI-shaped assistant message (content+tool_calls), soAgentOrchestratorruns identical logic regardless of provider.
Bringing your own HTTP client (PSR-18)
LlmService's constructor accepts a PSR-18 ClientInterface (plus PSR-17 request/stream
factories) — inject your own for connection pooling, retry/circuit-breaker middleware, or
tests that assert on the outgoing request without touching the network:
use Milpa\AiGateway\LlmService; $llm = new LlmService( apiKey: getenv('ANTHROPIC_API_KEY'), model: 'claude-3-5-sonnet-20241022', provider: 'anthropic', logger: $logger, // Psr\Log\LoggerInterface; optional httpClient: $yourPsr18Client, // Psr\Http\Client\ClientInterface; omit for Guzzle requestFactory: $yourFactory, // Psr\Http\Message\RequestFactoryInterface; optional streamFactory: $yourFactory, // Psr\Http\Message\StreamFactoryInterface; optional );
When httpClient is omitted, LlmService builds a Guzzle client with a 600s timeout
shared by both providers. That number used to be OpenAI-only-60s / Anthropic-only-600s (a
per-request Guzzle option on the Anthropic call, since Claude tool-use responses can run
long) — PSR-18's sendRequest() takes only a RequestInterface, with no per-call options
bag, so a per-provider timeout has no seam to hang off anymore. The default now simply
covers the slower case for both. Inject your own ClientInterface if you need the tighter
OpenAI-side timeout back.
Asking the provider for its context window: ProviderWindow
The number that governs compaction used to be a human assertion nothing verified — an app declared
a context window, and a declaration that was too large went unnoticed until the provider rejected
the prompt mid-run. ProviderWindow asks instead:
use Milpa\AiGateway\ProviderWindow; $window = new ProviderWindow('http://localhost:11438'); $window->tokens(); // 32768, or null when the provider did not say
Three things it does not do, each on purpose:
- It reads only the window the server ALLOCATED. Measured against
qwen3.8-27bon llama.cpp,GET /v1/modelsanswersdata[0].meta.n_ctx32768 next ton_ctx_train262144 — eight times apart. A reader that took the training window would believe it had 262k of room and blow up at 32k. When the training window is all a provider exposes, the answer isnull. - It never guesses and never raises. Unreachable, non-JSON, JSON without the field, a zero, a
negative, a float, a numeric string — every one of them resolves to
null. The caller asked for a better number, not for a new way to fail. - It asks once. The answer is memoised per instance, the silence included: asking is egress.
It tries GET /v1/models first — the OpenAI-compatible surface this gateway already drives, so any
base URL it can talk to answers there by construction — and falls back to llama.cpp's GET /props
(default_generation_settings.n_ctx), which carries the same figure but cannot be primary: the
measured payload declares "endpoint_props": false about itself.
The network is a seam, callable(string): ?string, so the whole reader is testable offline:
$window = new ProviderWindow('http://provider.test', fn (string $url): ?string => $recordedBody);
milpa/app-runtime composes this with whatever the app declared and keeps the smaller of the
two (greenhouse decisions/0233): a declaration is intent, the measurement is a ceiling that
cannot be exceeded.
What lives where
| Layer | Package | Owns |
|---|---|---|
| Contracts | milpa/tool-runtime |
LlmServiceInterface — the seam LlmService implements. |
| Tool execution | milpa/tool-runtime |
ToolRegistry, ToolContext, ToolResult, channel rendering — the pipeline McpClientService and AgentOrchestrator drive. |
| Gateway | milpa/ai-gateway (this package) |
The concrete LlmService (OpenAI + Anthropic, format translation both ways), McpClientService (registry facade), AgentOrchestrator (the ask-act loop), and ProviderWindow (the provider's allocated context window). |
| Your app | your host / plugins | API keys and secrets management, the PSR-3 logger you wire in, and any channel-specific glue (Telegram, web chat, CLI) around AgentOrchestrator::run(). |
Requirements
- PHP ≥ 8.3
milpa/core^0.6milpa/tool-runtime^0.5guzzlehttp/guzzle^7.10 — the default PSR-18 implementationLlmServicefalls back to when noClientInterfaceis injected (also bringsguzzlehttp/psr7, used as the default PSR-17 factory)psr/http-client,psr/http-factory,psr/http-message— the interfacesLlmService's constructor is typed againstpsr/log^3
Security note
LlmService can log provider request/response detail at debug level, including a slice of
the raw LLM response body — never enable that logging in production. See
SECURITY.md for the specifics.
Documentation
Full API reference: getmilpa.github.io/ai-gateway — generated straight from the source DocBlocks and dressed with the Milpa design system.
Contributing
Contributions are welcome — see CONTRIBUTING.md. Please report security issues via SECURITY.md, and note that this project follows a Code of Conduct.
License
Apache-2.0 © Rodrigo Vicente - TeamX Agency.
Milpa is designed, built, and maintained by Rodrigo Vicente - TeamX Agency.