milpa / ai-gateway
Dual-provider LLM gateway for the Milpa PHP framework: OpenAI + Anthropic chat-completions client with tool-call format translation, an agentic tool-use loop, and a ToolRegistry facade for LLM/MCP exposure.
Requires
- php: >=8.3
- guzzlehttp/guzzle: ^7.10
- guzzlehttp/psr7: ^2.7
- milpa/core: >=0.6.2 <1.0
- milpa/tool-runtime: >=0.18 <1.0
- psr/http-client: ^1.0
- psr/http-factory: ^1.1
- psr/http-message: ^2.0
- psr/log: ^3
Requires (Dev)
- friendsofphp/php-cs-fixer: ^3.65
- phpstan/phpdoc-parser: ^2.3
- phpstan/phpstan: ^2.1
- phpunit/phpunit: ^11.5
Suggests
None
Provides
None
Conflicts
None
Replaces
None
- dev-main
- v0.38.0
- v0.37.0
- v0.36.0
- v0.35.0
- v0.34.0
- v0.33.1
- v0.33.0
- v0.32.0
- v0.31.0
- v0.30.0
- v0.29.1
- v0.29.0
- v0.28.0
- v0.27.0
- v0.26.0
- v0.25.0
- v0.24.5
- v0.24.4
- v0.24.3
- v0.24.2
- v0.24.1
- v0.24.0
- v0.23.0
- v0.22.2
- v0.22.1
- v0.22.0
- v0.21.0
- v0.20.0
- v0.19.0
- v0.18.1
- v0.18.0
- v0.17.2
- v0.17.1
- v0.17.0
- v0.16.1
- v0.16.0
- v0.15.0
- v0.14.1
- v0.14.0
- v0.13.0
- v0.12.0
- v0.11.0
- v0.10.0
- v0.9.1
- v0.9.0
- v0.8.3
- v0.8.2
- v0.8.1
- v0.8.0
- v0.7.0
- v0.6.0
- v0.5.0
- v0.4.2
- v0.4.1
- v0.4.0
- v0.3.1
- v0.3.0
- v0.2.3
- v0.2.2
- v0.2.1
- v0.2.0
- v0.1.1
- v0.1.0
- dev-rodrigoteamx/local-template-thinking-0979
- dev-rodrigoteamx/local-reasoning-effort-0978
- dev-rodrigoteamx/ollama-max-effort-0968
- dev-rodrigoteamx/ollama-cloud-0955
- dev-rodrigoteamx/offered-tool-guard-0948
- dev-codex/current-offer-instructions-0867
- dev-feat/minimax-thinking-profile
- dev-rodrigoteamx/counted-input-budget
- dev-rodrigoteamx/context-yield
- dev-rodrigoteamx/context-projection
- dev-rodrigoteamx/native-output-budget
- dev-rodrigoteamx/diagnostic-output
- dev-rodrigoteamx/diagnostic-answer
- dev-rodrigoteamx/native-source-budget
- dev-rodrigoteamx/typed-termination
- dev-rodrigoteamx/structured-window
- dev-rodrigoteamx/enforce-withdrawal
- dev-rodrigoteamx/bounded-recovery
- dev-rodrigoteamx/lazy-tool-contracts
- dev-rodrigoteamx/output-contract
- dev-feat/stream-usage-token-cost
- dev-fix/flake-retry-twice
- dev-feat/context-self-healing
- dev-fix/provider-flake-retry
- dev-fix/null-tool-args
- dev-feat/intra-leg-budget
- dev-fix/notice-role
- dev-feat/forced-choice
- dev-feat/gateway-resilience
- dev-feat/reasoning-observer
- dev-feat/lazy-toolbox
This package is auto-updated.
Last update: 2026-09-23 02:58:44 UTC
README
Milpa AI Gateway
A dual-provider LLM gateway for the Milpa PHP framework — one client for OpenAI and Anthropic chat completions, translating each provider's tool-call wire format to and from a single shape, plus an agentic tool-use loop that drives a
milpa/tool-runtimeToolRegistry(resolve → validate → authorize → execute → audit) until the model is done.
milpa/ai-gateway is the LLM tier of Milpa: the piece that turns a milpa/tool-runtime
ToolRegistry into something a model can actually drive. LlmService implements
milpa/core's LlmServiceInterface seam against two concrete providers — OpenAI and
Anthropic — so callers write one message shape and one tool-call shape regardless of which
provider answers. AgentOrchestrator runs the loop every agent needs: ask the model, execute
whatever tools it asks for through the registry pipeline, feed the results back, repeat until
the model returns a final answer or a step budget runs out. No product coupling, no
Telegram/HTTP-specific code — those live in your host application.
Install
With AgentOrchestrator(lazyTools: true), describe_tool lists discoverable names and
descriptions. Describing a tool makes it callable with its full input schema on subsequent
requests. Undiscovered tools are not advertised as callable empty-object signatures. Discovery
does not authorize execution: the same tool registry and gates still judge every real call.
An empty catalogue exposes no discovery tool. The default full catalogue is unchanged.
composer require milpa/ai-gateway
Semantic recovery
An optional ProgressProbe owns the evidence and the recovery window. Its additive recovery
field reports pending, recovered, or exhausted. Pending recovery survives successful tool
calls and missing observations; only measured recovery clears it. Exhaustion returns
AgentOrchestrator::PROGRESS_STALLED with the producer's receipt before another model call.
Preparation can therefore lead to real work without making repeated bookkeeping count as work.
The total step budget remains in force. Probes omitting the field keep their previous behavior;
confirmation, debt and abandonment retain their exits, and abandonment does not reset recovery.
Quick example
Register a tool on a ToolRegistry (from milpa/tool-runtime), wrap it in McpClientService,
and hand both to AgentOrchestrator along with an LlmService:
use Milpa\AiGateway\AgentOrchestrator; use Milpa\AiGateway\LlmService; use Milpa\AiGateway\McpClientService; use Milpa\ToolRuntime\ToolRegistry; use Psr\Log\NullLogger; $registry = new ToolRegistry(new NullLogger()); $registry->register( 'get_time', 'Get the current time', [], fn () => ['time' => '12:00 PM'], ); $mcpClient = new McpClientService($registry); $llm = new LlmService(apiKey: getenv('OPENAI_API_KEY'), model: 'gpt-4o', provider: 'openai'); // LlmService talks HTTP through PSR-18 (Psr\Http\Client\ClientInterface), defaulting to a // Guzzle client when none is injected — see "Bringing your own HTTP client" below. $orchestrator = new AgentOrchestrator($llm, $mcpClient); echo $orchestrator->run('What time is it?'); // -> asks the model, the model requests `get_time`, AgentOrchestrator executes it through // the registry, feeds the result back, and returns the model's final answer.
Swap provider: 'anthropic' and a Claude model name (or let a claude model name in
$model select it automatically) to point the same call at Anthropic instead — LlmService
translates the tool list, the message history, and the tool-call response to and from
Anthropic's shape internally, so AgentOrchestrator and McpClientService never see a
provider-specific format.
Ollama Cloud uses its OpenAI-compatible endpoint through the same openai provider:
$llm = new LlmService( apiKey: getenv('OLLAMA_API_KEY'), model: 'glm-5.3-flash', provider: 'openai', baseUrl: 'https://ollama.com/v1', );
For this exact host, the gateway sends max_tokens, omits unsupported tool_choice, keeps
only standard OpenAI message fields, and preserves Bearer authentication and native tool-call
responses. Both https://ollama.com and https://ollama.com/v1 select this profile. Explicit
MiniMax thinking is rejected before transport because that request extension is provider-specific.
Use $llm->withOllamaReasoningEffort('low') (or medium / high / max) to send the documented
reasoning_effort option; omission preserves the model's default. The same explicit opt-in works
for another OpenAI-compatible endpoint, including a local llama.cpp server whose chat template
accepts the option. The gateway does not infer this capability for arbitrary hosts.
For a compatible local chat template, $llm->withOpenAiThinking(false) explicitly sends
chat_template_kwargs.enable_thinking=false; omission leaves the server's thinking behavior
unchanged.
Run termination
run(): string returns an answer or a terminal status, and propagates failures. After calling
the base loop, read $orchestrator->termination() for the producer's RunTermination observation:
use Milpa\AiGateway\RunEnd; try { $answer = $orchestrator->run('Inspect the app.'); } finally { $end = $orchestrator->termination(); // $end?->reason === RunEnd::FinalAnswer means the final-answer branch returned. // It does not mean the requested work is complete or approved. }
reason distinguishes final_answer, tool_refused, confirmation_required, blocked,
steps_exhausted, context_budget_exhausted, progress_stalled, house_debt, invalid_response, interrupted,
output_truncated, and failed. For progress_stalled, receipt carries the probe's original
receipt. toArray() exports both fields. A refusal and a final answer can have identical text;
the cause comes from the executed branch, never from matching that text.
The observation resets to null when each base run begins and is available after a return or
an exception escapes; the same exception object still propagates. Captured observations retain
their values when the instance is reused. The getter describes the latest base-loop run;
subclasses that replace run() must not present an earlier base-loop observation as a new one.
Concurrent or reentrant runs on the same orchestrator are not supported. Session questions,
completion evidence and authorization remain the host's responsibility.
The agent loop
AgentOrchestrator::run() alternates between two calls until the model is done or
$maxSteps (default 20) is reached:
- Ask —
LlmService::generateResponse()sends the running message history plus the registry's tool summaries (McpClientService::getToolSummaries()) to the provider and returns a single OpenAI-shaped assistant message. - Act — if that message carries
tool_calls, each one is executed viaMcpClientService::callTool(), which runs it through the fullToolRegistrypipeline (validate → authorize → execute → audit) under whateverToolContextwas set withsetToolContext(). The result — rendered through aRendererRegistrywhen one is configured, JSON otherwise — is fed back into the message history as atoolmessage, and the loop repeats.
setSystemPromptProjection(callable $projection) optionally derives the first system message
from the original system string and this request's exact tool summaries. The callback receives
(string $system, array $tools): string after lazy filtering and before context budgeting. It
changes only the outgoing copy: history, tool execution, the plan and progress notices retain
their existing paths. Withdrawal and restoration use a fresh projection of the same base;
no extra catalogue read or message accumulation occurs. Exceptions stop the run before the
model request. setSystemPromptProjection(null) restores the unchanged default path.
If a tool result requires confirmation or is blocked by policy, the loop stops immediately and returns that outcome instead of continuing — the caller (a chat handler, a CLI, a bot) is responsible for the confirm/cancel round trip on the next user turn.
AgentOrchestrator(..., outputTokens: 8192) declares one positive output limit for every loop
call, including existing recovery paths. Omitting it preserves the 4096-token default and the
previous input projection. With a known context, explicit output must be smaller than the
window; the input budget leaves room for it and refuses oversized protected input before egress.
The tool share uses LlmService::toolsForRequest(), the same provider-specific projection as
the transport: it counts wrappers and schema defaults, and excludes registry-only outputSchema
and version metadata. Messages, reasoning, tool-call pairs and the output reserve are preserved.
This uses the existing input estimator, not the provider's tokenizer. Unknown context still
transmits the finite limit but cannot establish available capacity. Truncation remains terminal
and never increases the budget or executes incomplete tools.
When an explicit context/output budget no longer fits after a completed step, the loop returns
CONTEXT_BUDGET_EXHAUSTED with the typed context_budget_exhausted cause. Its receipt records
estimatedInputTokens, contextTokens, outputTokens, inputLimitTokens, completedSteps,
and any pending progressReceipt / recovery. No additional request is sent, no tool is repeated,
and no answer judge runs. The caller may recompose its durable session and start a new leg within
its own finite total budget; the gateway does not resume automatically or certify useful progress.
An impossible initial projection still throws LengthException with failed; failures inside
provider recovery preserve their existing behavior. Unknown context or an omitted output budget
keep the previous behavior. A model quoting the sentinel is still an ordinary final answer.
An optional guard in a caller-supplied HTTP client can count the effective request, after
its transport transformations and before generation. The adapter owns the tokenizer, template,
provider identity and binding of the count to the exact body it would dispatch. If the count
exceeds contextTokens - outputTokens, it throws
InputBudgetExceededException(inputTokens, contextTokens, outputTokens, requestSha256).
If counting fails or cannot be bound to that body, it throws
InputBudgetUnavailableException(requestSha256, previous: $error) before dispatch.
This is an opt-in adapter contract; the gateway does not install a tokenizer, change defaults,
or treat an estimate as a count. A hash names a body; the adapter must actually verify it.
Both causes survive buffered/streamed OpenAI-compatible calls and Anthropic calls without a
transport retry. A counted overflow at a completed-step boundary returns
context_budget_exhausted only when the adapter's context/output limits match the loop's.
Its receipt uses source: counted_request, inputTokens and requestSha256 alongside the
limits, completed steps and pending recovery. It never labels the count estimatedInputTokens.
Initial refusals, unknown/mismatched limits and unavailable counts throw with failed and the
guard's receipt. Refusal during a degenerate-answer retry also throws: the original near-empty
answer cannot hide it. A genuine provider HTTP400 retains its existing healing path, but a
local guard refusal within that path does not trigger another generation retry.
Clients without such a guard retain the existing estimate, wire and behavior.
generateResponse(maxTokens: 4096) requires a positive output limit. It sends
max_completion_tokens to OpenAI-compatible endpoints and max_tokens to Anthropic.
Oversized structured tool exceptions use a bounded milpa.tool-failure-window/v1 projection.
Its outer ok:false describes the failed call; preview preserves the decoded error where it
fits, and partial plus omitted identify removed fields by JSON Pointer, character count and
SHA-256. original identifies the complete error by size and hash. The session recorder keeps
that original receipt; the projection does not write or replace it. Large object string fields
may be omitted, while arrays and scalar facts are never rewritten. If the structure still does
not fit, preview:null and a root omission report that it is unavailable. The same applies
when number spellings cannot survive PHP decoding and encoding unchanged. A preview is not a
complete receipt or authorization. Short errors, unsupported JSON and successful results keep
their existing paths; no context or per-result budget is increased.
When the provider reports output truncation (length / max_tokens), buffered and SSE
responses throw OutputTruncatedException before any tool call from that response can execute.
Anthropic's model_context_window_exceeded is also treated as truncation. The exception exposes
provider, stopReason and the requested maxTokens; it does not expose partial tool
arguments. Usage is still observed. There is no automatic retry, including during the guided
retry of a degenerate answer. Previously completed steps remain completed.
Stopping the loop before it acts: ToolCallGate
The loop above executes what the model asked for. ToolCallGate is the seam that lets somebody
else decide, before each call, whether it proceeds — and ToolCallRecorder is told after, with what
the tool answered. Both contracts live in milpa/tool-runtime (Milpa\ToolRuntime\Gate), where
every caller of tools already depends: a model is one caller, a governed door opened by a human is
another, and both ask the same question (greenhouse decisions/0225). This package keeps the two names
as interfaces that extend the base, so whoever implemented them here is still a gate or a recorder
wherever the base is asked for.
use Milpa\ToolRuntime\Gate\ToolCallGate; use Milpa\ToolRuntime\Gate\ToolCallRecorder; interface ToolCallGate { // The reason this call does not proceed, or null if it does. public function refuse(string $tool, array $arguments): ?string; } interface ToolCallRecorder { // Told after the call, with the rendered result and whether it succeeded. public function recorded(string $tool, array $arguments, string $result, bool $ok): void; }
refuse() runs before the call and recorded() after it. The two halves are not symmetric on
purpose: refusing is about INTENTION (what the caller is about to do), recording is about OUTCOME
(what happened). A refusal is not a tool error: McpClientService::callTool() throws
ToolCallRefusedException (which extends tool-runtime's ToolCallRefused) and the orchestrator
catches it apart, before any generic catch, and ends the turn. The gate is an interface here and
nothing else: this package brings no policy of its own — milpa/app-runtime implements it with the
session's permissions, the autonomy mode and the signatures.
The table: OptionTable
Since 0.5 the loop can operate against a table — the set of options the agent currently has in front of it — through one port with two questions that look alike and are not:
interface OptionTable { public function remove(string $option, string $code, ?string $message = null): void; public function removed(): array; // the PROJECTION: what the model sees public function wasRemoved(string $option): bool; // the FACT: did an authority declare it gone? }
The catalogue the model sees is re-derived every step from the registry minus removed() — it
is a projection, never a snapshot. And a refusal that removed the option no longer ends the run:
there is nothing left to route around, so the reason goes back to the model and the loop continues
with a different table. Both came out of measurement: a frozen catalogue made "did the agent re-read
the world?" unanswerable, and a run-ending refusal turned every denial into a shutdown (0 of 32
runs recovered).
SecondOpinionGate also stopped being quiet anywhere: an empty answer, a verdict-less answer and a
failed judge each say so with their own cause, a denial leaves a witness on stderr outside the
stream it writes to, and an ALLOW logs nothing — approving silently is correct; failing silently
was being counted as approval.
Provider translation
For an explicit finite JSON output shape, clone the client with a StructuredOutput:
$structured = $llm->withStructuredOutput(new \Milpa\AiGateway\StructuredOutput( 'diagnosis', ['required' => ['string'], 'configured' => ['string'], 'matches' => ['boolean']], )); $reply = $structured->generateResponse($prompt, $tools, $messages);
This preserves the client transport, observers, tools and output limit; $llm stays unchanged.
The clone sends response_format on buffered and streaming OpenAI-compatible calls. The provider
must support it; Anthropic is refused before HTTP, and HTTP errors never trigger a plain-text
fallback. The declaration admits at most 64 required scalar fields and no arbitrary schema,
constants or expected values. Replies remain unmodified: the caller must judge their format and
truth. Without the clone, requests retain their existing wire shape.
LlmService speaks one shape to its callers — OpenAI's messages / tool_calls — and
translates both directions for Anthropic:
- Outbound:
systemmessages become Anthropic's top-levelsystemparameter;toolrole messages becomeusermessages carrying atool_resultcontent block; an assistant message withtool_callsbecomestool_usecontent blocks. Tool summaries are reshaped from{name, description, inputSchema}to Anthropic's{name, description, input_schema}, with an emptypropertiesobject substituted where a tool declares none (Anthropic requires a non-empty schema object, not an empty array). - Inbound: Anthropic's
content: [{type: text, ...}, {type: tool_use, ...}]array is flattened back into a single OpenAI-shaped assistant message (content+tool_calls), soAgentOrchestratorruns identical logic regardless of provider.
Bringing your own HTTP client (PSR-18)
LlmService's constructor accepts a PSR-18 ClientInterface (plus PSR-17 request/stream
factories) — inject your own for connection pooling, retry/circuit-breaker middleware, or
tests that assert on the outgoing request without touching the network:
use Milpa\AiGateway\LlmService; $llm = new LlmService( apiKey: getenv('ANTHROPIC_API_KEY'), model: 'claude-3-5-sonnet-20241022', provider: 'anthropic', logger: $logger, // Psr\Log\LoggerInterface; optional httpClient: $yourPsr18Client, // Psr\Http\Client\ClientInterface; omit for Guzzle requestFactory: $yourFactory, // Psr\Http\Message\RequestFactoryInterface; optional streamFactory: $yourFactory, // Psr\Http\Message\StreamFactoryInterface; optional );
When httpClient is omitted, LlmService builds a Guzzle client with a 600s timeout
shared by both providers. That number used to be OpenAI-only-60s / Anthropic-only-600s (a
per-request Guzzle option on the Anthropic call, since Claude tool-use responses can run
long) — PSR-18's sendRequest() takes only a RequestInterface, with no per-call options
bag, so a per-provider timeout has no seam to hang off anymore. The default now simply
covers the slower case for both. Inject your own ClientInterface if you need the tighter
OpenAI-side timeout back.
Asking the provider for its context window: ProviderWindow
The number that governs compaction used to be a human assertion nothing verified — an app declared
a context window, and a declaration that was too large went unnoticed until the provider rejected
the prompt mid-run. ProviderWindow asks instead:
use Milpa\AiGateway\ProviderWindow; $window = new ProviderWindow('http://localhost:11438'); $window->tokens(); // 32768, or null when the provider did not say
Three things it does not do, each on purpose:
- It reads only the window the server ALLOCATED. Measured against
qwen3.8-27bon llama.cpp,GET /v1/modelsanswersdata[0].meta.n_ctx32768 next ton_ctx_train262144 — eight times apart. A reader that took the training window would believe it had 262k of room and blow up at 32k. When the training window is all a provider exposes, the answer isnull. - It never guesses and never raises. Unreachable, non-JSON, JSON without the field, a zero, a
negative, a float, a numeric string — every one of them resolves to
null. The caller asked for a better number, not for a new way to fail. - It asks once. The answer is memoised per instance, the silence included: asking is egress.
It tries GET /v1/models first — the OpenAI-compatible surface this gateway already drives, so any
base URL it can talk to answers there by construction — and falls back to llama.cpp's GET /props
(default_generation_settings.n_ctx), which carries the same figure but cannot be primary: the
measured payload declares "endpoint_props": false about itself.
The network is a seam, callable(string): ?string, so the whole reader is testable offline:
$window = new ProviderWindow('http://provider.test', fn (string $url): ?string => $recordedBody);
milpa/app-runtime composes this with whatever the app declared and keeps the smaller of the
two (greenhouse decisions/0233): a declaration is intent, the measurement is a ceiling that
cannot be exceeded.
What lives where
| Layer | Package | Owns |
|---|---|---|
| Contracts | milpa/tool-runtime |
LlmServiceInterface — the seam LlmService implements. |
| Tool execution | milpa/tool-runtime |
ToolRegistry, ToolContext, ToolResult, channel rendering — the pipeline McpClientService and AgentOrchestrator drive. |
| Gateway | milpa/ai-gateway (this package) |
The concrete LlmService (OpenAI + Anthropic, format translation both ways), McpClientService (registry facade), AgentOrchestrator (the ask-act loop), and ProviderWindow (the provider's allocated context window). |
| Your app | your host / plugins | API keys and secrets management, the PSR-3 logger you wire in, and any channel-specific glue (Telegram, web chat, CLI) around AgentOrchestrator::run(). |
Requirements
- PHP ≥ 8.3
milpa/core^0.6milpa/tool-runtime^0.5guzzlehttp/guzzle^7.10 — the default PSR-18 implementationLlmServicefalls back to when noClientInterfaceis injected (also bringsguzzlehttp/psr7, used as the default PSR-17 factory)psr/http-client,psr/http-factory,psr/http-message— the interfacesLlmService's constructor is typed againstpsr/log^3
Security note
LlmService can log provider request/response detail at debug level, including a slice of
the raw LLM response body — never enable that logging in production. See
SECURITY.md for the specifics.
Documentation
Full API reference: getmilpa.github.io/ai-gateway — generated straight from the source DocBlocks and dressed with the Milpa design system.
Contributing
Contributions are welcome — see CONTRIBUTING.md. Please report security issues via SECURITY.md, and note that this project follows a Code of Conduct.
License
Apache-2.0 © Rodrigo Vicente - TeamX Agency.
Milpa is designed, built, and maintained by Rodrigo Vicente - TeamX Agency.
Judging a final answer
A host can pass answerJudge: $judge to AgentOrchestrator. Its AnswerJudge::judge(string $candidate): AnswerVerdict receives the raw final candidate and returns accepted, rejected,
or indeterminate, bound to hash('sha256', $candidate), with a reason and producer evidence.
The host owns the criterion and evidence; the model cannot declare its own acceptance.
An accepted answer returns the exact candidate and terminates as final_answer, including after
a pending progress notice. Rejection terminates as answer_rejected; an unavailable judge or a
verdict for another candidate terminates as answer_indeterminate. RunTermination::toArray()
includes answerVerdict when present. The existing progress receipt remains unchanged: acceptance
does not manufacture growth, verify recorded work, grant authority, or add retries. Recovery
control markers, interruption, tool gates and the total budget retain their existing behavior.
With no judge, both the answer presentation and termination shape remain unchanged. See Greenhouse decision 0418 and evidence 0736 for the finite diagnostic consumer and falsification bank.
Producer-aware tool result limits
Each native tool invocation receives a ResultBudget through its ToolContext. The gateway derives the allowance from the declared context window using its existing rule (8000 characters for 32768 tokens; 6144 for 8192) and uses the same JSON encoder when delivering array results. A contextual handler can fit a complete result, including metadata, before returning it.
This does not raise the existing ceiling or change scope checks, consent, progress recovery, or history projection. Legacy strings keep their existing direct rendering; ToolResult objects keep their renderer path. Producers that ignore the optional budget retain the existing truncation behavior. The constraint lasts for one synchronous governed call.
Explicit MiniMax-M3 thinking
$llm->withMiniMaxThinking('disabled') returns a clone that sends
thinking: {"type": "disabled"} on the OpenAI-compatible MiniMax-M3 API.
adaptive is also supported; omission leaves provider defaults unchanged. Other
models, providers and modes are rejected before transport. This changes generation,
not the placement of reasoning text (reasoning_split), and does not change token
limits, messages, tool permissions or response parsing. Channel observers receive
the actual option. See Greenhouse evidence 0831.