@outputai/llm package is how you call LLMs from your steps and evaluators. It wraps the AI SDK and adds prompt files - version-controlled .prompt files that live alongside your code and define the provider, model, temperature, and prompt template in one place. Call arguments are the lists under Call arguments. AI SDK helpers and types (Output, tool, stepCountIs, ToolSet) are on the aiSdk namespace: import { generateText, aiSdk } from '@outputai/llm'.
Generate Functions
generateText is the primary function for LLM calls. Use the output parameter with aiSdk.Output.* helpers to control the response shape. When you need progress as text arrives, prefer generateTextWithStreaming in workflow steps: it uses streaming internally but returns a complete result and rejects on provider or transport errors. Use streamText when you need direct control over the stream:
Text Output
Generate unstructured text from a prompt file:steps.ts
result is a convenience alias for response.text.
Streaming
Complete result over streaming transport
generateTextWithStreaming behaves like generateText: await it to receive the complete response, including result, text, output, usage, finishReason, and cost. Internally it uses streaming transport and invokes onChunk as parts arrive.
This is the recommended streaming API for Output workflow steps. Provider, transport, and abort errors reject the returned promise, so the step fails and Temporal can apply its retry policy.
steps.ts
generateTextWithStreaming also supports structured output. Pass an aiSdk.Output.* specification and read the parsed value from result.output, just as with generateText.
Direct stream access
streamText remains available when you need to choose how the stream is consumed. It is not async: it returns a stream result synchronously, with textStream and fullStream iterables plus promise-based properties such as text, usage, and finishReason.
AI SDK streaming reports provider and transport failures through onError; consuming textStream does not reliably throw that original error. In a workflow step, capture the mapped error and throw it after consumption so Temporal records a failed activity instead of an empty successful result:
onError alone is not enough to fail the step. Output treats it as a fire-and-forget observer: exceptions and rejected promises from the callback are ignored to avoid a secondary stream failure. Capture the mapped error and throw it after consumption; awaiting a completion property can produce a generic no-output error instead of the original provider error.
Object Output
Generate a structured object matching a Zod schema. This is what you’ll use most in evaluators:evaluators.ts
output contains the typed object matching your schema.
Image Output
Generate images from a prompt file withgenerateImage. Image prompt files use plain instructions, not chat role tags like <system> or <user>. Keep plain text as the first meaningful body content so Output selects instruction mode:
prompts/nascar_race@v1.prompt
generateImage from a step:
steps.ts
result is a convenience alias for the first generated image (response.images[0]). The returned image exposes AI SDK image fields such as base64 and mediaType.
For image-to-image or edit flows, pass runtime image inputs with images and optionally mask. Output forwards these to the AI SDK prompt object:
Buffer, Uint8Array, ArrayBuffer, raw base64 strings, or { data, mediaType } objects. mask uses the same input shape and requires images.
generateImage does not upload generated images, download remote images, or normalize provider-specific values like size: "auto". Download or upload files in your workflow/client code, pass image bytes to images, and set concrete provider options in prompt front matter (size, n, aspectRatio, seed, providerOptions).Array Output
Generate an array of structured items:Choice Output
Select one value from a set of options:Agents
TheAgent class wraps AI SDK’s ToolLoopAgent with Output prompt files and the skills system. Use it when you need multi-step tool execution, conversation history, or a reusable agent instance with a fixed configuration. For single-shot LLM calls without tools, generateText is simpler.
Construction
The prompt file is loaded and rendered at construction time. Variables and tools are fixed at construction. Skills andmaxSteps come from the prompt file. The agent is ready to call generate(), generateWithStreaming(), or stream() immediately.
Each call seeds authored <user> blocks from the prompt. <system> blocks become instructions. Authored <assistant> blocks are dropped; use generateText when the prompt is a few-shot or prefilled thread.
steps.ts
generate()
Run the agent and return when complete:generateText: text, result (alias for text), output, usage, finishReason, toolCalls, etc.
Pass additional messages to extend the conversation. You can also pass abortSignal and toolChoice:
generateWithStreaming()
UsegenerateWithStreaming() when you want streaming progress and a complete result. It accepts the same messages, abortSignal, and toolChoice as generate(), plus onChunk:
generate(), the method returns the complete response, rejects on stream errors, and automatically appends messages to the configured message store. Prefer it over stream() in workflow steps unless you need direct access to the stream result.
stream()
Usestream() when you need direct access to the agent’s stream result. It accepts the same messages, abortSignal, and toolChoice as generate(), plus onChunk, onFinish, and onError:
streamText, the stream result provides textStream and fullStream iterables, plus promise-based properties (text, usage, finishReason) that resolve on completion. In a workflow step, capture and rethrow onError as shown so a failed stream cannot become an empty successful activity.
Structured Output
UseaiSdk.Output.object() with Agent to get typed responses:
steps.ts
Message Store
By default, Agent is stateless. Eachgenerate() / stream() call starts from the prompt seed (authored <user> blocks) plus this turn’s messages. Pass a messageStore to keep history across calls.
The store holds this turn’s caller messages plus the model reply. It does not persist the prompt seed. Reconstructing an agent is the same prompt (and variables) plus a hydrated store.
MessageStore is:
ModelMessage is an AI SDK type (aiSdk / ai). There is no built-in store. Implement the interface in memory for a single process, or with your database for durable history.
generate(), generateWithStreaming(), and stream() append messages to the message store. stream() stores in its wrapped onFinish when finishReason is not 'error'.When to Use Agent vs generateText
Start with
generateText. Move to Agent when you need conversation state or a reusable instance with a fixed configuration.
Response Object
generateText, generateTextWithStreaming, Agent.generate(), and Agent.generateWithStreaming() return the complete AI SDK response:
The
cost property is an LLM usage attribute:
usage. For example, reasoning is omitted when the model does not define separate reasoning pricing.
Direct streaming response shape. streamText and Agent.stream() return a different result type. Stream iterables (textStream, fullStream) provide real-time chunks, while scalar properties (text, usage, finishReason, etc.) are promises that resolve when the stream completes:
streamText / Agent.stream() onFinish receives the wrapped finish payload: result (alias for text), cost (null when pricing is missing), and merged sources (always an array).
Prompt Files
Instead of hardcoding model config and messages in your code, you write.prompt files that live in your workflow’s prompts/ folder. See the Prompts Guide for the full documentation.
prompts/generate_summary@v1.prompt
Configuration Options
Unknown top-level keys throw
Invalid prompt file. Snake_case aliases of known fields (max_tokens) include a suggestion (use "maxTokens"). Put provider-specific keys such as effort, reasoningEffort, and topP under providerOptions.
Providers
@outputai/llm ships built-in support for common AI SDK providers. The provider packages are peer dependencies with supported version ranges:
Legacy aliases
bedrock → amazon-bedrock and vertex → google-vertex are deprecated but still work; using them logs a deprecation warning.
Built-in provider instances are initialized lazily. Output creates the provider instance only when a prompt or API call first requests that provider, then reuses it for later calls.
Anthropic
ANTHROPIC_API_KEY environment variable.
OpenAI
OPENAI_API_KEY environment variable.
Azure OpenAI
AZURE_OPENAI_API_KEY, AZURE_OPENAI_ENDPOINT, and AZURE_OPENAI_API_VERSION.
Google Vertex AI
vertex is still accepted.
Amazon Bedrock
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION) or IAM role-based authentication. Set AWS_SESSION_TOKEN when using temporary credentials (e.g., from aws sts assume-role). The legacy alias bedrock is still accepted.
For cross-region inference, use the regional inference profile format: us.anthropic.claude-sonnet-4-20250514-v1:0.
Always set maxTokens in your Bedrock prompt files. Unlike the direct Anthropic provider (which auto-detects per-model limits), the Bedrock SDK has no client-side defaults and relies on server-side defaults that may be lower than the model’s capacity.
When using providerOptions, use the AI SDK bedrock namespace (not anthropic):
Custom Providers
UseregisterProvider when you want prompt files to reference an AI SDK provider that is not built in, or when you need a custom provider instance:
provider + model in the models.dev catalog. Built-in provider names match that catalog. A custom registerProvider name (such as vertex-anthropic above) will not match, so response.cost is null and a missing-cost warning is logged - register under a models.dev provider id if you need automatic pricing.
Prompt Caching
When a prompt sends the same large prefix on every call - a long system prompt, few-shot examples, a pasted reference document - you can cache that prefix so the provider skips reprocessing it. Cached input is about 90% cheaper and faster to first token. How you enable it depends on the provider.Anthropic
Anthropic caches only what you explicitly mark. Define acacheControl set in messageOptions and attach it - with options - to the block that ends your static prefix. Everything up to and including that block is cached and reused on the next call:
prompts/generate_summary@v1.prompt
<user> block - the part that changes each call - is re-billed at full price; the cached <system> prefix is charged at the much cheaper cache-read rate. For the 1-hour cache instead of the default 5 minutes, add ttl: 1h under cacheControl. A block can reference several sets (options="cached fast"), and a set can be reused across blocks. Bare options and names missing from messageOptions throw when the prompt loads.
Each set is a provider-namespaced providerOptions object - the same shape and namespace rules as prompt-file providerOptions. On Vertex with a Claude model, use the same anthropic namespace.
OpenAI
OpenAI caches automatically - there are no breakpoints to set, so themessageOptions mechanism above isn’t needed. Any prompt of 1024 tokens or longer is cached for you, with no markup. To improve hit rates across calls, set a stable promptCacheKey (and, on GPT-5.1+, extend retention) via providerOptions:
prompts/enrich_company@v1.prompt
Confirming a cache hit
Cache activity appears in the response usage and the cost event: the first call reports cache-creation tokens, and later calls within the TTL report cache-read tokens (cachedInputTokens), already priced at the cheaper rate in response.cost.
Anthropic caches only prefixes above a model-specific minimum - around 1,024 tokens for most Sonnet and Opus models, higher for some. Shorter prefixes are silently not cached, with no error. A request supports at most four cache breakpoints.
Provider Tools
Many providers offer built-in tools like web search. Configure them in YAML front matter:prompts/research@v1.prompt
tools.googleSearch({ mode: 'MODE_DYNAMIC', dynamicThreshold: 0.8 }) at the code level, but keeps your prompt self-contained.
YAML tools are merged with code-level tools, so you can combine provider tools (from YAML) with custom tools (from code). Code-level tools take precedence if names conflict.
For provider-specific tool options, see:
Tool Calling
Use tools withgenerateText to enable function calling:
Call arguments
variables accepts Liquid values, including nested objects and arrays. Use dot notation and Liquid loops to read structured values in the prompt template.
Text APIs
toolChoice, stopWhen, and prompt maxSteps apply only when tools exist. With tools, an explicit stopWhen takes precedence; otherwise Output uses aiSdk.stepCountIs(maxSteps) from the prompt (default 10).
generateImage
Agent
Constructor options are set onnew Agent(...). generate(), generateWithStreaming(), and stream() take an optional args object (or omit it). messages defaults to [].
Retries and Network Timeouts
Output always sets AI SDKmaxRetries to 0. In Output workflows, LLM calls usually run inside steps, and steps are Temporal activities. When a provider error fails the step, Temporal records the failed activity attempt and retries it according to the workflow’s retry policy.
generateText, generateTextWithStreaming, Agent.generate(), and Agent.generateWithStreaming() reject on failures. With direct streamText or Agent.stream() usage, capture the error in onError and throw it after consuming the stream, as shown in their direct-stream examples.
Built-in providers are initialized with a custom fetch that extends Undici’s headersTimeout and bodyTimeout to 15 minutes. This helps long-running LLM responses where the provider accepts the request but takes longer to return response headers or body chunks, for example reasoning-heavy calls. Active cancellation still works: if you pass abortSignal, or the AI SDK/provider aborts the request, that cancellation wins.
LLM call cost event
Each completed text generation call emits acost:llm:request event after the LLM responds and cost can be computed. For direct streams, the event is emitted when the stream finishes. You can observe it with the same hooks mechanism as error hooks: register a handler with on('cost:llm:request', handler) from @outputai/core/hooks in a hook file listed under outputai.hookFiles. The handler receives an event envelope whose payload field is the same LLM usage attribute exposed on response.cost. For payload details, see Cost Events.
loadPrompt
Load and render a prompt file without generating - useful for debugging:messages is empty and instructions contains the rendered body. For message-mode prompts, messages contains the parsed role blocks and instructions is null. See Prompt Body Modes.
PromptMessage.role is typed as 'system' | 'user' | 'assistant', matching the authored role blocks accepted in message mode.