Connecting a Real Model
Every agent in this chapter so far has run a model.example.ts: a few lines that follow a script. A real agent runs a real model, and FrontMCP comes with adapters for two providers, Anthropic and OpenAI. This lesson connects the triage agent to one, shows exactly what FrontMCP sends the provider and how it reads the answer, and covers what a real model changes: API keys, failures and retries, cost and waiting.
You will learn
- How to connect an agent to Anthropic or OpenAI, and which package each needs
- Where the API key comes from, and what happens when it's missing
- What FrontMCP sends the provider, and how it reads the reply
- What a failing provider does to a call, and how retries add up
- What a real model costs in requests, tokens and time
Choosing a provider
Replace the stand-in with provider, model and apiKey:
import { Agent, AgentContext, z } from "@frontmcp/sdk";
import { GetCustomer, GetTicket, SetPriority } from "./tools";
@Agent({
name: "triage",
description: "Triage a support ticket: read it, look up its customer, and set its priority. Pass the ticket id.",
systemInstructions: "You triage support tickets. …",
inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
tools: [GetTicket, GetCustomer, SetPriority],
llm: {
provider: "anthropic",
model: "claude-opus-5",
apiKey: { env: "ANTHROPIC_API_KEY" },
},
})
export class Triage extends AgentContext {}And install the provider's package next to @frontmcp/sdk. FrontMCP doesn't depend on either, and loads it the first time an agent asks the model:
yarn add @anthropic-ai/sdk # for provider: "anthropic"
yarn add openai # for provider: "openai"If the package is missing, the server still starts, and every call to the agent fails with The "@anthropic-ai/sdk" package is not installed., after about seven seconds of retries. That's a Node-only failure, so there's no example of it here.
llm option | Type | |
|---|---|---|
provider | "anthropic" or "openai" | Which adapter FrontMCP builds. |
model | string | The provider's model id, sent with every request. |
apiKey | { env: string } or string | Where the key comes from, below. |
baseUrl | string | Optional. Another address for the provider's API. |
maxTokens | number | Optional. The longest reply, in tokens. Anthropic's API requires one, and FrontMCP sends 4096 when you don't. |
temperature | number | Optional. Sent as is; leave it out for models that don't take it. |
provider is a shorthand: it builds an AnthropicAdapter or an OpenAIAdapter, both exported from @frontmcp/sdk. Build one yourself for what the shorthand doesn't cover, and pass it as llm: { adapter }:
import Anthropic from "@anthropic-ai/sdk";
import { AnthropicAdapter, OpenAIAdapter } from "@frontmcp/sdk";
// OpenAI's Responses API, instead of Chat Completions
new OpenAIAdapter({ model: "gpt-5", apiKey: process.env.OPENAI_API_KEY!, api: "responses" });
// Any service with an OpenAI-compatible API
new OpenAIAdapter({ model: "my-model", apiKey: process.env.GATEWAY_KEY!, baseUrl: "https://llm-gateway.internal/v1" });
// A client you've configured yourself, and fewer retries
new AnthropicAdapter({ model: "claude-opus-5", client: new Anthropic(), maxRetries: 1 });
For anything else, write an adapter: an object with completion(prompt, tools), like the stand-ins in this chapter, that calls your provider.
Where the key comes from
apiKey: { env: "ANTHROPIC_API_KEY" } reads the variable once, when the server starts. If it isn't set, the server doesn't start. In the Playground no variable is ever set, so the test builds a second server with a real provider, and watches it fail:
import { test, expect } from "@frontmcp/testing";
import { Agent, AgentContext, App, FrontMcpInstance } from "@frontmcp/sdk";
test("without the variable, the server doesn't start", async () => {
@Agent({
name: "triage",
llm: { provider: "anthropic", model: "claude-opus-5", apiKey: { env: "HELP_DESK_ANTHROPIC_KEY" } },
})
class Triage extends AgentContext {}
@App({ id: "help-desk", name: "Help Desk", agents: [Triage] })
class HelpDeskApp {}
const starting = FrontMcpInstance.createDirect({ info: { name: "help-desk", version: "1.0.0" }, apps: [HelpDeskApp] });
await expect(starting).rejects.toThrow("Environment variable HELP_DESK_ANTHROPIC_KEY is not set");
});Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
That's what you want: a server with a missing key fails at once, when it's deployed, not on the first call a user makes. apiKey also takes the key itself as a string, but then it's in your source code and everywhere your source code goes. Keep keys in the environment, or in your platform's secret store, which sets the environment for you.
The key is your server's, not the user's. Every call to the agent, from any client, is paid for by that key, so an agent is one more reason to authenticate clients and limit calls.
What FrontMCP sends the provider
An adapter turns FrontMCP's prompt into the provider's request, and the provider's reply back into { content, finishReason, toolCalls }. AnthropicAdapter takes an Anthropic client, so here it gets a stand-in that records each request and answers the way Anthropic's Messages API does. Open the Tests tab:
import { Agent, AgentContext, AnthropicAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import { anthropic } from "./anthropic.example";
@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string().describe("Ticket id, like T-1") } })
class GetTicket extends ToolContext {
async execute({ id }: { id: string }) {
return { id, title: "Cannot log in", status: "open" };
}
}
@Agent({
name: "triage",
description: "Triage a support ticket. Pass the ticket id.",
systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
tools: [GetTicket],
llm: { adapter: new AnthropicAdapter({ model: "claude-opus-5", client: anthropic }) },
})
export class Triage extends AgentContext {}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
OpenAIAdapter does the same for OpenAI's Chat Completions API. Its stand-in has chat.completions.create():
import { test, expect } from "@frontmcp/testing";
import { requests } from "./openai.example";
test("the instructions are the first message, and tools are functions", async ({ mcp }) => {
await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
expect(requests[0]).toMatchObject({
model: "gpt-5",
stream: false,
messages: [
{ role: "system", content: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why." },
{ role: "user", content: '{"ticketId":"T-1"}' },
],
tools: [{ type: "function", function: { name: "get_ticket", description: "Get a support ticket by its id." } }],
});
});
test("a tool result is a `tool` message", async ({ mcp }) => {
await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
expect(requests.at(-1).messages.slice(2)).toEqual([
{
role: "assistant",
content: null,
tool_calls: [{ id: "call_01", type: "function", function: { name: "get_ticket", arguments: '{"id":"T-1"}' } }],
},
{ role: "tool", content: '{"id":"T-1","title":"Cannot log in","status":"open"}', tool_call_id: "call_01" },
]);
});Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
| FrontMCP's prompt | Anthropic request | OpenAI request |
|---|---|---|
system | system | The first message, role: "system" |
A user message | role: "user" | role: "user" |
| The model's tool calls | tool_use blocks in an assistant message | tool_calls on an assistant message, arguments as a JSON string |
| A tool's result | A tool_result block in a user message | A tool message with tool_call_id |
| The agent's tools | tools, each with input_schema | tools, each { type: "function", function } |
Coming back, a reply with tool calls becomes finishReason: "tool_calls", and the loop runs them; anything else ends it. That includes a reply cut short by the token limit, which the loop treats as the answer.
When the provider fails
Providers fail: they rate-limit you, have outages, or reject your key. An adapter retries a failed request before it gives up, maxRetries times, 3 by default, waiting 1 second, then 2, then 4. When it gives up, the agent's call fails with the provider's message. This example keeps the waits short with maxRetries: 1:
import { Agent, AgentContext, OpenAIAdapter, z } from "@frontmcp/sdk";
import { openai } from "./openai.example";
@Agent({
name: "triage",
description: "Triage a support ticket. Pass the ticket id.",
inputSchema: { ticketId: z.string() },
llm: { adapter: new OpenAIAdapter({ model: "gpt-5", client: openai, maxRetries: 1 }) },
})
export class Triage extends AgentContext {}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
Retrying a rate limit is right: the second request worked. Retrying a wrong key isn't, because it will never work, but in 1.9.3 the adapter only gives up at once on a few messages, like invalid api key or unauthorized, and OpenAI's wording is Incorrect API key provided. With the default maxRetries: 3, a wrong key or a missing package costs every call seven seconds before it fails.
Two more things add up:
- Retries are per request. A loop that makes four requests can wait through retries four times, and all of it counts against the agent's
execution.timeout, 2 minutes by default. - The client waits. The client's model called one tool and is waiting for its answer. Keep
maxRetrieslow for agents that clients call directly, and setexecution.timeoutbelow what your clients wait for a tool call.
When the call fails, the development message above names the provider's error. In production, FrontMCP hides the message of a TOOL_EXECUTION_ERROR, so the client only learns that the call failed.
Cost and latency
Each turn of the loop is one request to the provider, and each request carries everything so far: the system instructions, every tool's schema, the input, and every tool call and result. So the input grows with every turn, and you pay for all of it every time. The adapter reports what each request used, as usage, and the loop drops it. To see it, wrap the adapter in one of your own, or override the agent's completion(), as the last challenge does:
import { Agent, AgentContext, AnthropicAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import type { AgentCompletion, AgentLlmAdapter } from "@frontmcp/sdk";
import { anthropic } from "./anthropic.example";
/** What each request to the model used, oldest first. */
export const usage: AgentCompletion["usage"][] = [];
function counted(inner: AgentLlmAdapter): AgentLlmAdapter {
return {
async completion(prompt, tools, options) {
const reply = await inner.completion(prompt, tools, options);
usage.push(reply.usage);
return reply;
},
};
}
@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string().describe("Ticket id, like T-1") } })
class GetTicket extends ToolContext {
async execute({ id }: { id: string }) {
return { id, title: "Cannot log in", status: "open", body: "Since this morning nobody on our team can log in." };
}
}
@Agent({
name: "triage",
description: "Triage a support ticket. Pass the ticket id.",
systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
tools: [GetTicket],
llm: { adapter: counted(new AnthropicAdapter({ model: "claude-opus-5", client: anthropic, maxTokens: 12 })) },
})
export class Triage extends AgentContext {}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
maxTokens caps each reply, and this example sets it far too low on purpose. The provider stopped mid-sentence, and FrontMCP's loop took the half sentence as the answer: it ends at any reply that isn't a tool call, including one that ran out of tokens, and says nothing about it. Leave maxTokens at the default, or raise it; lower it only when you know how long answers are.
What makes an agent cost more, and take longer:
- The number of requests. Every turn is a request, and the client waits for all of them.
execution.maxIterationscaps it, 10 by default. - What each request repeats. Instructions, tool descriptions and schemas are sent every turn. Give an agent the tools its task needs, not every tool you have.
- What tools return. Every result stays in the conversation until the end. A tool an agent uses should return what the model needs, not the whole record: the same advice as for shaping tool results, and it matters more here, because the result is paid for again on every later turn.
- Where it runs. An agent spends your key on behalf of whoever calls it. Put agents behind authentication and rate limits.
The other direction: connectOpenAI() and friends
FrontMCP also exports connectOpenAI(), connectClaude(), connectLangChain() and connectVercelAI(). Despite the names, they don't connect an agent to a model. They go the other way: they connect your code to a FrontMCP server, in-process, and hand you its tools in the shape that provider's API or library expects. They're for when your own program runs the model and the loop, and the FrontMCP server is only the tools. Running FrontMCP Anywhere shows them.
| You want | Use |
|---|---|
| Your MCP server to run a model loop of its own, as a tool clients call | @Agent with llm |
| Your own program to run the model loop, with a FrontMCP server's tools | connectClaude(), connectOpenAI(), … |
Recap
llm: { provider: "anthropic" | "openai", model, apiKey: { env } }connects an agent to a real model. Install@anthropic-ai/sdkoropenai; FrontMCP loads it on the first request.apiKey: { env }is read when the server starts, and a missing variable stops the server from starting. Keep keys out of your source.AnthropicAdapterandOpenAIAdapterare whatproviderbuilds. Build one yourself forapi: "responses",baseUrl, your ownclient, ormaxRetries.- Adapters retry a failed request 3 times, waiting 1, 2 and 4 seconds, including failures that can't succeed, like a wrong key. A call that still fails is a
TOOL_EXECUTION_ERROR. - Every turn is a request carrying the whole conversation. Wrap the adapter, or override the agent's
completion(), to seeusage. A reply cut off atmaxTokensbecomes the answer, silently. connectOpenAI()and friends connect your own code to a FrontMCP server; they aren't model adapters.
Try some challenges
Each challenge runs hidden checks against your code. Edit the code, then press Check.
Challenge 1 of 2
Fail fast when the provider is busy
This agent is called directly by clients, and the provider is rate-limiting every request. With maxRetries: 2, each call waits three seconds before it fails. Make a rate-limited call fail at once, after a single request.
import { Agent, AgentContext, OpenAIAdapter, z } from "@frontmcp/sdk";
import { openai } from "./openai.example";
@Agent({
name: "triage",
description: "Triage a support ticket. Pass the ticket id.",
inputSchema: { ticketId: z.string() },
llm: { adapter: new OpenAIAdapter({ model: "gpt-5", client: openai, maxRetries: 2 }) },
})
export class Triage extends AgentContext {}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.