feat(llm): provider-executed tool pass-through

Add a `providerExecuted: boolean` flag to `tool-call` and `tool-result`
events plus the persisted `ToolResultPart`. When set, the tool runtime
skips client dispatch (the provider already executed the tool) and folds
both events into the assistant message so the next round's history
carries the call + result for context.

Anthropic: decode `server_tool_use` blocks and the three server tool
result block types (`web_search_tool_result`, `code_execution_tool_result`,
`web_fetch_tool_result`) into `tool-call` / `tool-result` events with
`providerExecuted: true`. Round-trip the same parts back into the
provider when the assistant message is replayed in subsequent requests.
Result block error payloads (`*_tool_result_error`) surface as
`result.type === "error"`.

OpenAI Responses: decode hosted tool items emitted via
`response.output_item.done` (`web_search_call`, `file_search_call`,
`code_interpreter_call`, `computer_use_call`, `image_generation_call`,
`mcp_call`, `local_shell_call`) as `tool-call` + `tool-result` pairs
with `providerExecuted: true`. Each tool's input fields are pulled out
explicitly; the full item is passed through as the result payload so
consumers can read outputs / sources / status without re-decoding.

Tool runtime: extend the dispatch decision so provider-executed
tool-calls bypass the handler lookup, and tool-result events with
`providerExecuted: true` are appended to the assistant content for
round-trip rather than being treated as a separate tool message.

Tests: 7 new deterministic fixtures cover Anthropic decode (success +
error result + round-trip + unknown server tool name), OpenAI Responses
decode (web_search_call, code_interpreter_call), and tool-runtime
skip-dispatch.

AGENTS.md updates the runtime section to describe pass-through behavior
and notes the transport-agnostic design that keeps a future WebSocket
adapter (e.g. OpenAI Codex backend) as a sibling rather than a core
rewrite.
This commit is contained in:
Kit Langton
2026-04-26 12:33:58 -04:00
parent b5ca62d1ea
commit e1c6bf92fb
8 changed files with 530 additions and 9 deletions
+11 -2
View File
@@ -47,6 +47,8 @@ Adapters should stay boring and typed:
- `toHttp` creates the `HttpClientRequest`.
- `parse` decodes provider chunks into `LLMEvent`s. The shared `ProviderShared.sse` helper handles SSE framing, chunk decoding, and stateful chunk-to-event raising; adapters supply `decodeChunk` and a `process` callback that produces events.
The transport is HTTP + SSE today; the `LLMEvent` stream contract is intentionally transport-agnostic. When a provider ships a non-HTTP transport (OpenAI's WebSocket-based Codex backend, hypothetical bidirectional streaming APIs), it should land as a sibling adapter with a `toWs` (or analogous) producer + a `parse` that reads frames from that transport — not by leaking transport details into core types.
### Patches
Patches are the forcing function for provider/model quirks. If a behavior is not universal enough for common IR, keep it as a named patch with a trace entry. Good examples:
@@ -133,7 +135,13 @@ Errors must be expressed as `ToolFailure`. The runtime catches it and emits a `t
- Input failed the `parameters` Schema.
- The handler returned a `ToolFailure`.
Provider-defined tools (e.g. OpenAI built-in `web_search`) should go directly into `request.tools` without a runtime entry. The runtime currently raises `tool-error` for unknown names; if you need pass-through, file an issue.
Provider-defined / hosted tools (e.g. Anthropic `web_search` / `code_execution` / `web_fetch`, OpenAI Responses `web_search_call` / `file_search_call` / `code_interpreter_call` / `mcp_call` / `local_shell_call` / `image_generation_call` / `computer_use_call`) pass through the runtime untouched:
- Adapters surface the model's call as a `tool-call` event with `providerExecuted: true`, and the provider's result as a matching `tool-result` event with `providerExecuted: true`.
- The runtime detects `providerExecuted` on `tool-call` and **skips client dispatch** — no handler is invoked and no `tool-error` is raised for "unknown tool". The provider already executed it.
- Both events are appended to the assistant message in `assistantContent` so the next round's history carries the call + result for context. Anthropic encodes them back as `server_tool_use` + `web_search_tool_result` (or `code_execution_tool_result` / `web_fetch_tool_result`) blocks; OpenAI Responses callers typically use `previous_response_id` instead of resending hosted-tool items.
Add provider-defined tools to `request.tools` (no runtime entry needed). The matching adapter must know how to lower the tool definition into the provider-native shape; right now Anthropic accepts `web_search` / `code_execution` / `web_fetch` and OpenAI Responses accepts the hosted tool names listed above.
### Recording Tests
@@ -193,7 +201,8 @@ Do not blanket re-record an entire test file when adding one cassette. `RECORD=t
- [ ] Build a `Provider.Model` -> `LLM.ModelRef` bridge for OpenCode, including protocol selection, base URLs, headers, limits, capabilities, native provider metadata, and OpenAI-compatible provider family detection.
- [ ] Build a `session.llm` -> `LLM.request(...)` bridge for system prompts, message history, tools, tool choice, generation options, reasoning variants, cache hints, and attachments.
- [x] Add a typed `ToolRuntime` that drives the tool loop with Schema-typed parameters/success per tool, single-`ToolFailure` error channel, and `maxSteps`/`stopWhen` controls. Provider-defined tool pass-through is still TODO.
- [x] Add a typed `ToolRuntime` that drives the tool loop with Schema-typed parameters/success per tool, single-`ToolFailure` error channel, and `maxSteps`/`stopWhen` controls.
- [x] Provider-defined tool pass-through: `providerExecuted` flag on `tool-call`/`tool-result` events; Anthropic `server_tool_use` / `web_search_tool_result` / `code_execution_tool_result` / `web_fetch_tool_result` round-trip; OpenAI Responses hosted-tool items decoded as `tool-call` + `tool-result` pairs; runtime skips client dispatch when `providerExecuted: true`.
- [ ] Keep auth and deployment concerns in the OpenCode bridge where possible: Bedrock credentials/region/profile, Vertex project/location/token, Azure deployment/API version, and Gateway/OpenRouter routing headers.
- [ ] Keep initial OpenCode integration behind a local flag/path until request payload parity and stream event parity are proven against the existing `session/llm.test.ts` cases.
+117 -4
View File
@@ -48,6 +48,35 @@ const AnthropicToolUseBlock = Schema.Struct({
})
type AnthropicToolUseBlock = Schema.Schema.Type<typeof AnthropicToolUseBlock>
const AnthropicServerToolUseBlock = Schema.Struct({
type: Schema.Literal("server_tool_use"),
id: Schema.String,
name: Schema.String,
input: Schema.Unknown,
cache_control: Schema.optional(AnthropicCacheControl),
})
type AnthropicServerToolUseBlock = Schema.Schema.Type<typeof AnthropicServerToolUseBlock>
// Server tool result blocks: web_search_tool_result, code_execution_tool_result,
// and web_fetch_tool_result. The provider executes the tool and inlines the
// structured result into the assistant turn — there is no client tool_result
// round-trip. We round-trip the structured `content` payload as opaque JSON so
// the next request can echo it back when continuing the conversation.
const AnthropicServerToolResultType = Schema.Literals([
"web_search_tool_result",
"code_execution_tool_result",
"web_fetch_tool_result",
])
type AnthropicServerToolResultType = Schema.Schema.Type<typeof AnthropicServerToolResultType>
const AnthropicServerToolResultBlock = Schema.Struct({
type: AnthropicServerToolResultType,
tool_use_id: Schema.String,
content: Schema.Unknown,
cache_control: Schema.optional(AnthropicCacheControl),
})
type AnthropicServerToolResultBlock = Schema.Schema.Type<typeof AnthropicServerToolResultBlock>
const AnthropicToolResultBlock = Schema.Struct({
type: Schema.Literal("tool_result"),
tool_use_id: Schema.String,
@@ -57,7 +86,13 @@ const AnthropicToolResultBlock = Schema.Struct({
})
const AnthropicUserBlock = Schema.Union([AnthropicTextBlock, AnthropicToolResultBlock])
const AnthropicAssistantBlock = Schema.Union([AnthropicTextBlock, AnthropicThinkingBlock, AnthropicToolUseBlock])
const AnthropicAssistantBlock = Schema.Union([
AnthropicTextBlock,
AnthropicThinkingBlock,
AnthropicToolUseBlock,
AnthropicServerToolUseBlock,
AnthropicServerToolResultBlock,
])
type AnthropicAssistantBlock = Schema.Schema.Type<typeof AnthropicAssistantBlock>
type AnthropicToolResultBlock = Schema.Schema.Type<typeof AnthropicToolResultBlock>
@@ -118,6 +153,11 @@ const AnthropicStreamBlock = Schema.Struct({
text: Schema.optional(Schema.String),
thinking: Schema.optional(Schema.String),
input: Schema.optional(Schema.Unknown),
// *_tool_result blocks arrive whole as content_block_start (no streaming
// delta) with the structured payload in `content` and the originating
// server_tool_use id in `tool_use_id`.
tool_use_id: Schema.optional(Schema.String),
content: Schema.optional(Schema.Unknown),
})
const AnthropicStreamDelta = Schema.Struct({
@@ -145,6 +185,7 @@ interface ToolAccumulator {
readonly id: string
readonly name: string
readonly input: string
readonly providerExecuted: boolean
}
interface ParserState {
@@ -200,6 +241,29 @@ const lowerToolCall = (part: ToolCallPart): AnthropicToolUseBlock => ({
input: part.input,
})
const lowerServerToolCall = (part: ToolCallPart): AnthropicServerToolUseBlock => ({
type: "server_tool_use",
id: part.id,
name: part.name,
input: part.input,
})
// Server tool result blocks are typed by name. Anthropic ships three today;
// extend this list when new server tools land. The block content is the
// structured payload returned by the provider, which we round-trip as-is.
const serverToolResultType = (name: string): AnthropicServerToolResultType | undefined => {
if (name === "web_search") return "web_search_tool_result"
if (name === "code_execution") return "code_execution_tool_result"
if (name === "web_fetch") return "web_fetch_tool_result"
return undefined
}
const lowerServerToolResult = Effect.fn("AnthropicMessages.lowerServerToolResult")(function* (part: ToolResultPart) {
const wireType = serverToolResultType(part.name)
if (!wireType) return yield* invalid(`Anthropic Messages does not know how to round-trip server tool result for ${part.name}`)
return { type: wireType, tool_use_id: part.id, content: part.result.value } satisfies AnthropicServerToolResultBlock
})
const lowerMessages = Effect.fn("AnthropicMessages.lowerMessages")(function* (request: LLMRequest) {
const messages: AnthropicMessage[] = []
@@ -226,7 +290,11 @@ const lowerMessages = Effect.fn("AnthropicMessages.lowerMessages")(function* (re
continue
}
if (part.type === "tool-call") {
content.push(lowerToolCall(part))
content.push(part.providerExecuted ? lowerServerToolCall(part) : lowerToolCall(part))
continue
}
if (part.type === "tool-result" && part.providerExecuted) {
content.push(yield* lowerServerToolResult(part))
continue
}
return yield* invalid(`Anthropic Messages assistant messages only support text, reasoning, and tool-call content for now`)
@@ -337,9 +405,44 @@ const finishToolCall = (tool: ToolAccumulator | undefined) =>
tool.input || "{}",
`Invalid JSON input for Anthropic Messages tool call ${tool.name}`,
)
return [{ type: "tool-call" as const, id: tool.id, name: tool.name, input }]
const event: LLMEvent = tool.providerExecuted
? { type: "tool-call", id: tool.id, name: tool.name, input, providerExecuted: true }
: { type: "tool-call", id: tool.id, name: tool.name, input }
return [event]
})
// Server tool result blocks come whole in `content_block_start` (no streaming
// delta sequence). We convert the payload to a `tool-result` event with
// `providerExecuted: true`. The runtime appends it to the assistant message
// for round-trip; downstream consumers can inspect `result.value` for the
// structured payload.
const SERVER_TOOL_RESULT_NAMES: Record<AnthropicServerToolResultType, string> = {
web_search_tool_result: "web_search",
code_execution_tool_result: "code_execution",
web_fetch_tool_result: "web_fetch",
}
const isServerToolResultType = (type: string): type is AnthropicServerToolResultType =>
type in SERVER_TOOL_RESULT_NAMES
const serverToolResultEvent = (block: NonNullable<AnthropicChunk["content_block"]>): LLMEvent | undefined => {
if (!block.type || !isServerToolResultType(block.type)) return undefined
const errorPayload =
typeof block.content === "object" && block.content !== null && "type" in block.content
? String((block.content as Record<string, unknown>).type)
: ""
const isError = errorPayload.endsWith("_tool_result_error")
return {
type: "tool-result",
id: block.tool_use_id ?? "",
name: SERVER_TOOL_RESULT_NAMES[block.type],
result: isError
? { type: "error", value: block.content }
: { type: "json", value: block.content },
providerExecuted: true,
}
}
const processChunk = (state: ParserState, chunk: AnthropicChunk) =>
Effect.gen(function* () {
if (chunk.type === "message_start") {
@@ -347,7 +450,11 @@ const processChunk = (state: ParserState, chunk: AnthropicChunk) =>
return [usage ? { ...state, usage: mergeUsage(state.usage, usage) } : state, []] as const
}
if (chunk.type === "content_block_start" && chunk.index !== undefined && chunk.content_block?.type === "tool_use") {
if (
chunk.type === "content_block_start" &&
chunk.index !== undefined &&
(chunk.content_block?.type === "tool_use" || chunk.content_block?.type === "server_tool_use")
) {
return [{
...state,
tools: {
@@ -356,6 +463,7 @@ const processChunk = (state: ParserState, chunk: AnthropicChunk) =>
id: chunk.content_block.id ?? String(chunk.index),
name: chunk.content_block.name ?? "",
input: "",
providerExecuted: chunk.content_block.type === "server_tool_use",
},
},
}, []] as const
@@ -369,6 +477,11 @@ const processChunk = (state: ParserState, chunk: AnthropicChunk) =>
return [state, [{ type: "reasoning-delta", text: chunk.content_block.thinking }]] as const
}
if (chunk.type === "content_block_start" && chunk.content_block) {
const event = serverToolResultEvent(chunk.content_block)
if (event) return [state, [event]] as const
}
if (chunk.type === "content_block_delta" && chunk.delta?.type === "text_delta" && chunk.delta.text) {
return [state, [{ type: "text-delta", text: chunk.delta.text }]] as const
}
@@ -94,7 +94,22 @@ const OpenAIResponsesStreamItem = Schema.Struct({
call_id: Schema.optional(Schema.String),
name: Schema.optional(Schema.String),
arguments: Schema.optional(Schema.String),
// Hosted (provider-executed) tool fields. Each hosted tool item carries its
// own subset of these — we capture them generically so we can surface the
// call's typed input portion and round-trip the full result payload without
// hand-rolling a per-tool schema.
status: Schema.optional(Schema.String),
action: Schema.optional(Schema.Unknown),
queries: Schema.optional(Schema.Unknown),
results: Schema.optional(Schema.Unknown),
code: Schema.optional(Schema.String),
container_id: Schema.optional(Schema.String),
outputs: Schema.optional(Schema.Unknown),
server_label: Schema.optional(Schema.String),
output: Schema.optional(Schema.Unknown),
error: Schema.optional(Schema.Unknown),
})
type OpenAIResponsesStreamItem = Schema.Schema.Type<typeof OpenAIResponsesStreamItem>
const OpenAIResponsesChunk = Schema.Struct({
type: Schema.String,
@@ -275,6 +290,57 @@ const finishToolCall = (tools: Record<string, ToolAccumulator>, item: NonNullabl
return [{ type: "tool-call" as const, id: item.call_id, name: item.name, input }]
})
// Hosted tool items (provider-executed) ship their typed input + status + result
// fields all in one item. We expose them as a `tool-call` + `tool-result` pair
// so consumers can treat them uniformly with client tools, only differentiated
// by `providerExecuted: true`.
//
// item.type → tool name. Each entry is the OpenAI Responses item type that
// represents a hosted (provider-executed) tool call.
const HOSTED_TOOL_NAMES: Record<string, string> = {
web_search_call: "web_search",
web_search_preview_call: "web_search_preview",
file_search_call: "file_search",
code_interpreter_call: "code_interpreter",
computer_use_call: "computer_use",
image_generation_call: "image_generation",
mcp_call: "mcp",
local_shell_call: "local_shell",
}
const isHostedToolItem = (item: OpenAIResponsesStreamItem): item is OpenAIResponsesStreamItem & { id: string } =>
item.type in HOSTED_TOOL_NAMES && typeof item.id === "string" && item.id.length > 0
// Pick the input fields the model actually populated when invoking the tool.
// The shape is tool-specific. Keep this list explicit so each tool's input is
// reviewable at a glance — fall back to `{}` for tools we haven't typed yet.
const hostedToolInput = (item: OpenAIResponsesStreamItem): unknown => {
if (item.type === "web_search_call" || item.type === "web_search_preview_call") return item.action ?? {}
if (item.type === "file_search_call") return { queries: item.queries ?? [] }
if (item.type === "code_interpreter_call") return { code: item.code, container_id: item.container_id }
if (item.type === "computer_use_call") return item.action ?? {}
if (item.type === "local_shell_call") return item.action ?? {}
if (item.type === "mcp_call") return { server_label: item.server_label, name: item.name, arguments: item.arguments }
return {}
}
// Round-trip the full item as the structured result so consumers can extract
// outputs / sources / status without re-decoding.
const hostedToolResult = (item: OpenAIResponsesStreamItem) => {
const isError = typeof item.error !== "undefined" && item.error !== null
return isError
? ({ type: "error" as const, value: item.error })
: ({ type: "json" as const, value: item })
}
const hostedToolEvents = (item: OpenAIResponsesStreamItem & { id: string }): ReadonlyArray<LLMEvent> => {
const name = HOSTED_TOOL_NAMES[item.type]!
return [
{ type: "tool-call", id: item.id, name, input: hostedToolInput(item), providerExecuted: true },
{ type: "tool-result", id: item.id, name, result: hostedToolResult(item), providerExecuted: true },
]
}
const processChunk = (state: ParserState, chunk: OpenAIResponsesChunk) =>
Effect.gen(function* () {
if (chunk.type === "response.output_text.delta" && chunk.delta) {
@@ -306,6 +372,10 @@ const processChunk = (state: ParserState, chunk: OpenAIResponsesChunk) =>
return [state, events] as const
}
if (chunk.type === "response.output_item.done" && chunk.item && isHostedToolItem(chunk.item)) {
return [state, hostedToolEvents(chunk.item)] as const
}
if (chunk.type === "response.completed" || chunk.type === "response.incomplete") {
return [state, [{ type: "request-finish" as const, reason: mapFinishReason(chunk), usage: mapUsage(chunk.response?.usage) }]] as const
}
+3
View File
@@ -121,6 +121,7 @@ export const ToolResultPart = Schema.Struct({
id: Schema.String,
name: Schema.String,
result: ToolResultValue,
providerExecuted: Schema.optional(Schema.Boolean),
metadata: Schema.optional(Schema.Record(Schema.String, Schema.Unknown)),
}).annotate({ identifier: "LLM.Content.ToolResult" })
export type ToolResultPart = Schema.Schema.Type<typeof ToolResultPart>
@@ -262,6 +263,7 @@ export const ToolCall = Schema.Struct({
id: Schema.String,
name: Schema.String,
input: Schema.Unknown,
providerExecuted: Schema.optional(Schema.Boolean),
}).annotate({ identifier: "LLM.Event.ToolCall" })
export type ToolCall = Schema.Schema.Type<typeof ToolCall>
@@ -270,6 +272,7 @@ export const ToolResult = Schema.Struct({
id: Schema.String,
name: Schema.String,
result: ToolResultValue,
providerExecuted: Schema.optional(Schema.Boolean),
}).annotate({ identifier: "LLM.Event.ToolResult" })
export type ToolResult = Schema.Schema.Type<typeof ToolResult>
+24 -2
View File
@@ -10,6 +10,7 @@ import {
type LLMEvent,
LLMRequest,
type ToolCallPart,
type ToolResultPart,
type ToolResultValue,
} from "./schema"
import { ToolFailure } from "./schema"
@@ -127,9 +128,30 @@ const accumulate = (state: StepState, event: LLMEvent) => {
return
}
if (event.type === "tool-call") {
const part: ToolCallPart = { type: "tool-call", id: event.id, name: event.name, input: event.input }
const part: ToolCallPart = {
type: "tool-call",
id: event.id,
name: event.name,
input: event.input,
providerExecuted: event.providerExecuted,
}
state.assistantContent.push(part)
// Provider-executed tools are dispatched by the provider; the runtime must
// not invoke a client handler. The matching `tool-result` event arrives
// later in the same stream and is folded into `assistantContent` so the
// next round's message history carries it.
if (!event.providerExecuted) state.toolCalls.push(part)
return
}
if (event.type === "tool-result" && event.providerExecuted) {
const part: ToolResultPart = {
type: "tool-result",
id: event.id,
name: event.name,
result: event.result,
providerExecuted: true,
}
state.assistantContent.push(part)
state.toolCalls.push(part)
return
}
if (event.type === "request-finish") {
@@ -163,6 +163,174 @@ describe("Anthropic Messages adapter", () => {
}),
)
it.effect("decodes server_tool_use + web_search_tool_result as provider-executed events", () =>
Effect.gen(function* () {
const body = sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{ type: "content_block_start", index: 0, content_block: { type: "server_tool_use", id: "srvtoolu_abc", name: "web_search" } },
{ type: "content_block_delta", index: 0, delta: { type: "input_json_delta", partial_json: '{"query":"effect 4"}' } },
{ type: "content_block_stop", index: 0 },
{
type: "content_block_start",
index: 1,
content_block: {
type: "web_search_tool_result",
tool_use_id: "srvtoolu_abc",
content: [{ type: "web_search_result", url: "https://example.com", title: "Example" }],
},
},
{ type: "content_block_stop", index: 1 },
{ type: "content_block_start", index: 2, content_block: { type: "text", text: "" } },
{ type: "content_block_delta", index: 2, delta: { type: "text_delta", text: "Found it." } },
{ type: "content_block_stop", index: 2 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 8 } },
)
const response = yield* client({ adapters: [AnthropicMessages.adapter] })
.generate(
LLM.request({
...request,
tools: [{ name: "web_search", description: "Web search", inputSchema: { type: "object" } }],
}),
)
.pipe(Effect.provide(fixedResponse(body)))
const toolCall = response.events.find((event) => event.type === "tool-call")
expect(toolCall).toEqual({
type: "tool-call",
id: "srvtoolu_abc",
name: "web_search",
input: { query: "effect 4" },
providerExecuted: true,
})
const toolResult = response.events.find((event) => event.type === "tool-result")
expect(toolResult).toEqual({
type: "tool-result",
id: "srvtoolu_abc",
name: "web_search",
result: { type: "json", value: [{ type: "web_search_result", url: "https://example.com", title: "Example" }] },
providerExecuted: true,
})
expect(LLM.outputText(response)).toBe("Found it.")
expect(response.events.at(-1)).toMatchObject({ type: "request-finish", reason: "stop" })
}),
)
it.effect("decodes web_search_tool_result_error as provider-executed error result", () =>
Effect.gen(function* () {
const body = sseEvents(
{ type: "message_start", message: { usage: { input_tokens: 5 } } },
{ type: "content_block_start", index: 0, content_block: { type: "server_tool_use", id: "srvtoolu_x", name: "web_search" } },
{ type: "content_block_delta", index: 0, delta: { type: "input_json_delta", partial_json: '{"query":"q"}' } },
{ type: "content_block_stop", index: 0 },
{
type: "content_block_start",
index: 1,
content_block: {
type: "web_search_tool_result",
tool_use_id: "srvtoolu_x",
content: { type: "web_search_tool_result_error", error_code: "max_uses_exceeded" },
},
},
{ type: "content_block_stop", index: 1 },
{ type: "message_delta", delta: { stop_reason: "end_turn" }, usage: { output_tokens: 1 } },
)
const response = yield* client({ adapters: [AnthropicMessages.adapter] })
.generate(
LLM.request({
...request,
tools: [{ name: "web_search", description: "Web search", inputSchema: { type: "object" } }],
}),
)
.pipe(Effect.provide(fixedResponse(body)))
const toolResult = response.events.find((event) => event.type === "tool-result")
expect(toolResult).toMatchObject({
type: "tool-result",
id: "srvtoolu_x",
name: "web_search",
result: { type: "error" },
providerExecuted: true,
})
}),
)
it.effect("round-trips provider-executed assistant content into server tool blocks", () =>
Effect.gen(function* () {
const prepared = yield* client({ adapters: [AnthropicMessages.adapter] }).prepare(
LLM.request({
id: "req_round_trip",
model,
messages: [
LLM.user("Search for something."),
LLM.assistant([
{
type: "tool-call",
id: "srvtoolu_abc",
name: "web_search",
input: { query: "effect 4" },
providerExecuted: true,
},
{
type: "tool-result",
id: "srvtoolu_abc",
name: "web_search",
result: { type: "json", value: [{ url: "https://example.com" }] },
providerExecuted: true,
},
{ type: "text", text: "Found it." },
]),
LLM.user("Thanks."),
],
}),
)
expect(prepared.target).toMatchObject({
messages: [
{ role: "user", content: [{ type: "text", text: "Search for something." }] },
{
role: "assistant",
content: [
{ type: "server_tool_use", id: "srvtoolu_abc", name: "web_search", input: { query: "effect 4" } },
{
type: "web_search_tool_result",
tool_use_id: "srvtoolu_abc",
content: [{ url: "https://example.com" }],
},
{ type: "text", text: "Found it." },
],
},
{ role: "user", content: [{ type: "text", text: "Thanks." }] },
],
})
}),
)
it.effect("rejects round-trip for unknown server tool names", () =>
Effect.gen(function* () {
const error = yield* client({ adapters: [AnthropicMessages.adapter] })
.prepare(
LLM.request({
id: "req_unknown_server_tool",
model,
messages: [
LLM.assistant([
{
type: "tool-result",
id: "srvtoolu_abc",
name: "future_server_tool",
result: { type: "json", value: {} },
providerExecuted: true,
},
]),
],
}),
)
.pipe(Effect.flip)
expect(error.message).toContain("future_server_tool")
}),
)
it.effect("rejects unsupported user media content", () =>
Effect.gen(function* () {
const error = yield* client({ adapters: [AnthropicMessages.adapter] })
@@ -158,6 +158,80 @@ describe("OpenAI Responses adapter", () => {
}),
)
it.effect("decodes web_search_call as provider-executed tool-call + tool-result", () =>
Effect.gen(function* () {
const item = {
type: "web_search_call",
id: "ws_1",
status: "completed",
action: { type: "search", query: "effect 4" },
}
const body = sseEvents(
{ type: "response.output_item.added", item },
{ type: "response.output_item.done", item },
{ type: "response.completed", response: { usage: { input_tokens: 5, output_tokens: 1 } } },
)
const response = yield* client({ adapters: [OpenAIResponses.adapter] })
.generate(request)
.pipe(Effect.provide(fixedResponse(body)))
const callsAndResults = response.events.filter((event) => event.type === "tool-call" || event.type === "tool-result")
expect(callsAndResults).toEqual([
{
type: "tool-call",
id: "ws_1",
name: "web_search",
input: { type: "search", query: "effect 4" },
providerExecuted: true,
},
{
type: "tool-result",
id: "ws_1",
name: "web_search",
result: { type: "json", value: item },
providerExecuted: true,
},
])
}),
)
it.effect("decodes code_interpreter_call as provider-executed events with code input", () =>
Effect.gen(function* () {
const item = {
type: "code_interpreter_call",
id: "ci_1",
status: "completed",
code: "print(1+1)",
container_id: "cnt_xyz",
outputs: [{ type: "logs", logs: "2\n" }],
}
const body = sseEvents(
{ type: "response.output_item.done", item },
{ type: "response.completed", response: { usage: { input_tokens: 5, output_tokens: 1 } } },
)
const response = yield* client({ adapters: [OpenAIResponses.adapter] })
.generate(request)
.pipe(Effect.provide(fixedResponse(body)))
const toolCall = response.events.find((event) => event.type === "tool-call")
expect(toolCall).toEqual({
type: "tool-call",
id: "ci_1",
name: "code_interpreter",
input: { code: "print(1+1)", container_id: "cnt_xyz" },
providerExecuted: true,
})
const toolResult = response.events.find((event) => event.type === "tool-result")
expect(toolResult).toEqual({
type: "tool-result",
id: "ci_1",
name: "code_interpreter",
result: { type: "json", value: item },
providerExecuted: true,
})
}),
)
it.effect("rejects unsupported user media content", () =>
Effect.gen(function* () {
const error = yield* client({ adapters: [OpenAIResponses.adapter] })
+63 -1
View File
@@ -1,7 +1,8 @@
import { describe, expect } from "bun:test"
import { Effect, Layer, Schema, Stream } from "effect"
import { LLM, LLMEvent } from "../src"
import { client } from "../src/adapter"
import { client, type LLMClient } from "../src/adapter"
import { RequestExecutor } from "../src/executor"
import { OpenAIChat } from "../src/provider/openai-chat"
import { tool, ToolFailure } from "../src/tool"
import { ToolRuntime } from "../src/tool-runtime"
@@ -184,6 +185,67 @@ describe("ToolRuntime", () => {
}),
)
it.effect("does not dispatch provider-executed tool calls", () =>
Effect.gen(function* () {
// Stub client emits a provider-executed tool-call followed by its
// tool-result and a stop. The runtime must not dispatch a handler (no
// tool-error for unknown name) and must not loop (no second stream).
let streams = 0
const stub: LLMClient = {
prepare: () => Effect.die("not used"),
generate: () => Effect.die("not used"),
stream: () => {
streams++
return Stream.fromIterable<LLMEvent>([
{ type: "request-start", id: "req_1", model: baseRequest.model },
{
type: "tool-call",
id: "srvtoolu_abc",
name: "web_search",
input: { query: "x" },
providerExecuted: true,
},
{
type: "tool-result",
id: "srvtoolu_abc",
name: "web_search",
result: { type: "json", value: { results: [] } },
providerExecuted: true,
},
{ type: "text-delta", text: "Done." },
{ type: "request-finish", reason: "stop" },
])
},
}
// The runtime's stream type carries `RequestExecutor.Service` because
// adapters use it. Our stub never executes HTTP, but the type still
// demands the service — provide a noop so the test compiles.
const noopExecutor = Layer.succeed(RequestExecutor.Service, {
execute: () => Effect.die("stub client never executes HTTP"),
})
const events = Array.from(
yield* ToolRuntime.run(stub, { request: baseRequest, tools: {} }).pipe(
Stream.runCollect,
Effect.provide(noopExecutor),
),
)
expect(streams).toBe(1)
expect(events.find(LLMEvent.guards["tool-error"])).toBeUndefined()
expect(events.filter(LLMEvent.guards["tool-call"])).toEqual([
{
type: "tool-call",
id: "srvtoolu_abc",
name: "web_search",
input: { query: "x" },
providerExecuted: true,
},
])
expect(LLM.outputText({ events })).toBe("Done.")
}),
)
it.effect("dispatches multiple tool calls in one step concurrently", () =>
Effect.gen(function* () {
const llm = client({ adapters: [OpenAIChat.adapter] })