# Connecting a Real Model

> How to connect a FrontMCP agent to Anthropic or OpenAI, where the API key comes from, what FrontMCP sends the provider and how it reads the reply, what a failing provider does to a call, and what a real model costs in requests, tokens and time.

Source: https://frontmcp.dev/learn/connecting-a-real-model

Every agent in this chapter so far has run a `model.example.ts`: a few lines that follow a script. A real agent runs a real model, and FrontMCP comes with adapters for two providers, Anthropic and OpenAI. This lesson connects the triage agent to one, shows exactly what FrontMCP sends the provider and how it reads the answer, and covers what a real model changes: API keys, failures and retries, cost and waiting.

**You will learn**
- How to connect an agent to Anthropic or OpenAI, and which package each needs
- Where the API key comes from, and what happens when it's missing
- What FrontMCP sends the provider, and how it reads the reply
- What a failing provider does to a call, and how retries add up
- What a real model costs in requests, tokens and time

> **Note**
The Playground can't reach a model provider: it runs in your browser, with no network and no API key, and the providers' packages aren't in it. So the configuration in this lesson is plain code, to run on your own machine. The examples that do run use FrontMCP's real adapters, with a stand-in for the provider's client, the way the earlier lessons used a stand-in for the model.

## Choosing a provider

Replace the stand-in with `provider`, `model` and `apiKey`:

```ts triage.agent.ts
import { Agent, AgentContext, z } from "@frontmcp/sdk";
import { GetCustomer, GetTicket, SetPriority } from "./tools";

@Agent({
  name: "triage",
  description: "Triage a support ticket: read it, look up its customer, and set its priority. Pass the ticket id.",
  systemInstructions: "You triage support tickets. …",
  inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
  tools: [GetTicket, GetCustomer, SetPriority],
  llm: {
    provider: "anthropic",
    model: "claude-opus-5",
    apiKey: { env: "ANTHROPIC_API_KEY" },
  },
})
export class Triage extends AgentContext {}
```

And install the provider's package next to `@frontmcp/sdk`. FrontMCP doesn't depend on either, and loads it the first time an agent asks the model:

```bash
yarn add @anthropic-ai/sdk   # for provider: "anthropic"
yarn add openai              # for provider: "openai"
```

If the package is missing, the server still starts, and every call to the agent fails with `The "@anthropic-ai/sdk" package is not installed.`, after about seven seconds of [retries](#when-the-provider-fails). That's a Node-only failure, so there's no example of it here.

| `llm` option | Type | |
| --- | --- | --- |
| `provider` | `"anthropic"` or `"openai"` | Which adapter FrontMCP builds. |
| `model` | `string` | The provider's model id, sent with every request. |
| `apiKey` | `{ env: string }` or `string` | Where the key comes from, [below](#where-the-key-comes-from). |
| `baseUrl` | `string` | Optional. Another address for the provider's API. |
| `maxTokens` | `number` | Optional. The longest reply, in tokens. Anthropic's API requires one, and FrontMCP sends `4096` when you don't. |
| `temperature` | `number` | Optional. Sent as is; leave it out for models that don't take it. |

`provider` is a shorthand: it builds an `AnthropicAdapter` or an `OpenAIAdapter`, both exported from `@frontmcp/sdk`. Build one yourself for what the shorthand doesn't cover, and pass it as `llm: { adapter }`:

```ts
import Anthropic from "@anthropic-ai/sdk";
import { AnthropicAdapter, OpenAIAdapter } from "@frontmcp/sdk";

// OpenAI's Responses API, instead of Chat Completions
new OpenAIAdapter({ model: "gpt-5", apiKey: process.env.OPENAI_API_KEY!, api: "responses" });

// Any service with an OpenAI-compatible API
new OpenAIAdapter({ model: "my-model", apiKey: process.env.GATEWAY_KEY!, baseUrl: "https://llm-gateway.internal/v1" });

// A client you've configured yourself, and fewer retries
new AnthropicAdapter({ model: "claude-opus-5", client: new Anthropic(), maxRetries: 1 });
```

For anything else, write an adapter: an object with `completion(prompt, tools)`, like the stand-ins in this chapter, that calls your provider.

## Where the key comes from

`apiKey: { env: "ANTHROPIC_API_KEY" }` reads the variable once, when the server starts. If it isn't set, the server doesn't start. In the Playground no variable is ever set, so the test builds a second server with a real provider, and watches it fail:

```ts key.test.ts active
import { test, expect } from "@frontmcp/testing";
import { Agent, AgentContext, App, FrontMcpInstance } from "@frontmcp/sdk";

test("without the variable, the server doesn't start", async () => {
  @Agent({
    name: "triage",
    llm: { provider: "anthropic", model: "claude-opus-5", apiKey: { env: "HELP_DESK_ANTHROPIC_KEY" } },
  })
  class Triage extends AgentContext {}

  @App({ id: "help-desk", name: "Help Desk", agents: [Triage] })
  class HelpDeskApp {}

  const starting = FrontMcpInstance.createDirect({ info: { name: "help-desk", version: "1.0.0" }, apps: [HelpDeskApp] });
  await expect(starting).rejects.toThrow("Environment variable HELP_DESK_ANTHROPIC_KEY is not set");
});
```

```ts triage.agent.ts
import { Agent, AgentContext, z } from "@frontmcp/sdk";
import { model } from "./model.example";

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  inputSchema: { ticketId: z.string() },
  llm: { adapter: model },
})
export class Triage extends AgentContext {}
```

```ts model.example.ts
// Stands in for a real model, so this example's own server starts.
import type { AgentLlmAdapter } from "@frontmcp/sdk";

export const model: AgentLlmAdapter = {
  async completion() {
    return { content: "T-1 is high priority.", finishReason: "stop" };
  },
};
```

That's what you want: a server with a missing key fails at once, when it's deployed, not on the first call a user makes. `apiKey` also takes the key itself as a string, but then it's in your source code and everywhere your source code goes. Keep keys in the environment, or in your platform's secret store, which sets the environment for you.

The key is your server's, not the user's. Every call to the agent, from any client, is paid for by that key, so an agent is one more reason to [authenticate clients](https://frontmcp.dev/learn/authenticating-clients) and [limit calls](https://frontmcp.dev/learn/limiting-calls).

## What FrontMCP sends the provider

An adapter turns FrontMCP's prompt into the provider's request, and the provider's reply back into `{ content, finishReason, toolCalls }`. `AnthropicAdapter` takes an Anthropic client, so here it gets a stand-in that records each request and answers the way Anthropic's Messages API does. Open the **Tests** tab:

```ts triage.agent.ts active
import { Agent, AgentContext, AnthropicAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import { anthropic } from "./anthropic.example";

@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string().describe("Ticket id, like T-1") } })
class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
  inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
  tools: [GetTicket],
  llm: { adapter: new AnthropicAdapter({ model: "claude-opus-5", client: anthropic }) },
})
export class Triage extends AgentContext {}
```

```ts anthropic.example.ts
// Stands in for the client from @anthropic-ai/sdk, which the Playground can't
// use: there's no network and no API key. It has the one method FrontMCP calls,
// messages.create(), records every request, and answers in the Messages API's
// shape: first with a tool call, then with text. A real server passes
// `new Anthropic()`, or uses `provider: "anthropic"`.

/** Every request "Anthropic" received, oldest first. */
export const requests: any[] = [];

export const anthropic = {
  messages: {
    async create(params: any): Promise<any> {
      requests.push(structuredClone(params));
      if (params.messages.length === 1) {
        return {
          content: [{ type: "tool_use", id: "toolu_01", name: "get_ticket", input: { id: "T-1" } }],
          stop_reason: "tool_use",
          usage: { input_tokens: 612, output_tokens: 48 },
        };
      }
      return {
        content: [{ type: "text", text: "High: the customer can't log in." }],
        stop_reason: "end_turn",
        usage: { input_tokens: 703, output_tokens: 12 },
      };
    },
  },
};
```

```ts anthropic.test.ts
import { test, expect } from "@frontmcp/testing";
import { requests } from "./anthropic.example";

test("the first request: instructions, input and tools", async ({ mcp }) => {
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(requests[0]).toEqual({
    model: "claude-opus-5",
    max_tokens: 4096,
    stream: false,
    system: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
    messages: [{ role: "user", content: '{"ticketId":"T-1"}' }],
    tools: [
      {
        name: "get_ticket",
        description: "Get a support ticket by its id.",
        input_schema: {
          $schema: "https://json-schema.org/draft/2020-12/schema",
          type: "object",
          properties: { id: { type: "string", description: "Ticket id, like T-1" } },
          required: ["id"],
        },
      },
    ],
  });
});

test("the second request adds the tool call and its result", async ({ mcp }) => {
  const before = requests.length;
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(requests.length - before).toBe(2);
  expect(requests.at(-1).messages).toEqual([
    { role: "user", content: '{"ticketId":"T-1"}' },
    { role: "assistant", content: [{ type: "tool_use", id: "toolu_01", name: "get_ticket", input: { id: "T-1" } }] },
    {
      role: "user",
      content: [{ type: "tool_result", tool_use_id: "toolu_01", content: '{"id":"T-1","title":"Cannot log in","status":"open"}' }],
    },
  ]);
});

test("the text reply is the agent's answer", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result.json()).toEqual({ response: "High: the customer can't log in." });
});
```

`OpenAIAdapter` does the same for OpenAI's Chat Completions API. Its stand-in has `chat.completions.create()`:

```ts openai.test.ts active
import { test, expect } from "@frontmcp/testing";
import { requests } from "./openai.example";

test("the instructions are the first message, and tools are functions", async ({ mcp }) => {
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(requests[0]).toMatchObject({
    model: "gpt-5",
    stream: false,
    messages: [
      { role: "system", content: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why." },
      { role: "user", content: '{"ticketId":"T-1"}' },
    ],
    tools: [{ type: "function", function: { name: "get_ticket", description: "Get a support ticket by its id." } }],
  });
});

test("a tool result is a `tool` message", async ({ mcp }) => {
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(requests.at(-1).messages.slice(2)).toEqual([
    {
      role: "assistant",
      content: null,
      tool_calls: [{ id: "call_01", type: "function", function: { name: "get_ticket", arguments: '{"id":"T-1"}' } }],
    },
    { role: "tool", content: '{"id":"T-1","title":"Cannot log in","status":"open"}', tool_call_id: "call_01" },
  ]);
});
```

```ts triage.agent.ts
import { Agent, AgentContext, OpenAIAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import { openai } from "./openai.example";

@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string().describe("Ticket id, like T-1") } })
class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
  inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
  tools: [GetTicket],
  llm: { adapter: new OpenAIAdapter({ model: "gpt-5", client: openai }) },
})
export class Triage extends AgentContext {}
```

```ts openai.example.ts
// Stands in for the client from the openai package, which the Playground can't
// use. It records every request and answers in the Chat Completions shape.

/** Every request "OpenAI" received, oldest first. */
export const requests: any[] = [];

export const openai = {
  chat: {
    completions: {
      async create(params: any): Promise<any> {
        requests.push(structuredClone(params));
        const message = params.messages.some((m: any) => m.role === "tool")
          ? { role: "assistant", content: "High: the customer can't log in." }
          : { role: "assistant", content: null, tool_calls: [{ id: "call_01", type: "function", function: { name: "get_ticket", arguments: '{"id":"T-1"}' } }] };
        return {
          choices: [{ message, finish_reason: message.content ? "stop" : "tool_calls" }],
          usage: { prompt_tokens: 540, completion_tokens: 20, total_tokens: 560 },
        };
      },
    },
  },
  responses: {
    async create(): Promise<any> {
      throw new Error("This example uses Chat Completions.");
    },
  },
};
```

| FrontMCP's prompt | Anthropic request | OpenAI request |
| --- | --- | --- |
| `system` | `system` | The first message, `role: "system"` |
| A `user` message | `role: "user"` | `role: "user"` |
| The model's tool calls | `tool_use` blocks in an `assistant` message | `tool_calls` on an `assistant` message, `arguments` as a JSON string |
| A tool's result | A `tool_result` block in a `user` message | A `tool` message with `tool_call_id` |
| The agent's tools | `tools`, each with `input_schema` | `tools`, each `{ type: "function", function }` |

Coming back, a reply with tool calls becomes `finishReason: "tool_calls"`, and the loop runs them; anything else ends it. That includes a reply cut short by the token limit, which [the loop treats as the answer](#cost-and-latency).

## When the provider fails

Providers fail: they rate-limit you, have outages, or reject your key. An adapter retries a failed request before it gives up, `maxRetries` times, 3 by default, waiting 1 second, then 2, then 4. When it gives up, the agent's call fails with the provider's message. This example keeps the waits short with `maxRetries: 1`:

```ts triage.agent.ts active
import { Agent, AgentContext, OpenAIAdapter, z } from "@frontmcp/sdk";
import { openai } from "./openai.example";

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  inputSchema: { ticketId: z.string() },
  llm: { adapter: new OpenAIAdapter({ model: "gpt-5", client: openai, maxRetries: 1 }) },
})
export class Triage extends AgentContext {}
```

```ts openai.example.ts
// Stands in for the openai client. The first request for T-1 is rate-limited,
// and every request for T-2 is refused, as if the key were wrong.

export const received: string[] = [];

function fail(status: number, message: string) {
  return Object.assign(new Error(message), { status });
}

export const openai = {
  chat: {
    completions: {
      async create(params: any): Promise<any> {
        const ticket = JSON.parse(params.messages.at(-1).content).ticketId;
        received.push(ticket);
        if (ticket === "T-2") throw fail(401, "401 Incorrect API key provided: sk-proj-****. You can find your API key at https://platform.openai.com/account/api-keys.");
        if (received.filter((t) => t === ticket).length === 1) throw fail(429, "429 Rate limit reached for gpt-5");
        return { choices: [{ message: { role: "assistant", content: `${ticket} is high priority.` }, finish_reason: "stop" }] };
      },
    },
  },
  responses: {
    async create(): Promise<any> {
      throw new Error("This example uses Chat Completions.");
    },
  },
};
```

```ts retries.test.ts
import { test, expect } from "@frontmcp/testing";
import { received } from "./openai.example";

test("a rate-limited request is tried again a second later", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result.json()).toEqual({ response: "T-1 is high priority." });
  expect(received.filter((t) => t === "T-1")).toHaveLength(2);
  expect(result.durationMs).toBeGreaterThanOrEqual(1_000);
});

test("🚩 a wrong key is tried again too, and then fails the call", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-2" });
  expect(received.filter((t) => t === "T-2")).toHaveLength(2);
  expect(result).toBeError("TOOL_EXECUTION_ERROR");
  expect(result.text()).toMatch(/^Tool "invoke_triage" execution failed: 401 Incorrect API key provided/);
});
```

Retrying a rate limit is right: the second request worked. Retrying a wrong key isn't, because it will never work, but in 1.9.3 the adapter only gives up at once on a few messages, like `invalid api key` or `unauthorized`, and OpenAI's wording is `Incorrect API key provided`. With the default `maxRetries: 3`, a wrong key or a missing package costs every call seven seconds before it fails.

Two more things add up:

- **Retries are per request.** A loop that makes four requests can wait through retries four times, and all of it counts against the agent's [`execution.timeout`](https://frontmcp.dev/learn/giving-an-agent-tools#limiting-the-loop), 2 minutes by default.
- **The client waits.** The client's model called one tool and is waiting for its answer. Keep `maxRetries` low for agents that clients call directly, and set `execution.timeout` below what your clients wait for a tool call.

When the call fails, the development message above names the provider's error. In production, FrontMCP hides the message of a `TOOL_EXECUTION_ERROR`, so the client only learns that the call failed.

## Cost and latency

Each turn of the loop is one request to the provider, and each request carries everything so far: the system instructions, every tool's schema, the input, and every tool call and result. So the input grows with every turn, and you pay for all of it every time. The adapter reports what each request used, as `usage`, and the loop drops it. To see it, wrap the adapter in one of your own, or override the agent's [`completion()`](https://frontmcp.dev/learn/giving-an-agent-tools#each-request-and-each-tool-call), as the [last challenge](#try-some-challenges) does:

```ts triage.agent.ts active
import { Agent, AgentContext, AnthropicAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import type { AgentCompletion, AgentLlmAdapter } from "@frontmcp/sdk";
import { anthropic } from "./anthropic.example";

/** What each request to the model used, oldest first. */
export const usage: AgentCompletion["usage"][] = [];

function counted(inner: AgentLlmAdapter): AgentLlmAdapter {
  return {
    async completion(prompt, tools, options) {
      const reply = await inner.completion(prompt, tools, options);
      usage.push(reply.usage);
      return reply;
    },
  };
}

@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string().describe("Ticket id, like T-1") } })
class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open", body: "Since this morning nobody on our team can log in." };
  }
}

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
  inputSchema: { ticketId: z.string().describe("Ticket id, like T-1") },
  tools: [GetTicket],
  llm: { adapter: counted(new AnthropicAdapter({ model: "claude-opus-5", client: anthropic, maxTokens: 12 })) },
})
export class Triage extends AgentContext {}
```

```ts anthropic.example.ts
// Stands in for the Anthropic client. It counts a token as four characters of
// the request, and cuts its answer off at max_tokens, as a real model does.

export const anthropic = {
  messages: {
    async create(params: any): Promise<any> {
      const input_tokens = Math.ceil(JSON.stringify(params).length / 4);
      if (params.messages.length === 1) {
        return {
          content: [{ type: "tool_use", id: "toolu_01", name: "get_ticket", input: { id: "T-1" } }],
          stop_reason: "tool_use",
          usage: { input_tokens, output_tokens: 40 },
        };
      }
      const words = "High priority: nobody at the customer can log in, so their whole team is blocked.".split(" ");
      const cut = words.length > params.max_tokens;
      return {
        content: [{ type: "text", text: words.slice(0, params.max_tokens).join(" ") }],
        stop_reason: cut ? "max_tokens" : "end_turn",
        usage: { input_tokens, output_tokens: Math.min(words.length, params.max_tokens) },
      };
    },
  },
};
```

```ts cost.test.ts
import { test, expect } from "@frontmcp/testing";
import { usage } from "./triage.agent";

test("each request reports its usage, and the second is bigger", async ({ mcp }) => {
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  const [first, second] = usage;
  expect(first).toEqual({ promptTokens: expect.any(Number), completionTokens: 40, totalTokens: expect.any(Number) });
  expect(second!.promptTokens).toBeGreaterThan(first!.promptTokens);
});

test("🚩 an answer cut off at `maxTokens` is still the answer", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result).toBeSuccessful();
  expect(result.json()).toEqual({ response: "High priority: nobody at the customer can log in, so their whole" });
});
```

`maxTokens` caps each reply, and this example sets it far too low on purpose. The provider stopped mid-sentence, and FrontMCP's loop took the half sentence as the answer: it ends at any reply that isn't a tool call, including one that ran out of tokens, and says nothing about it. Leave `maxTokens` at the default, or raise it; lower it only when you know how long answers are.

What makes an agent cost more, and take longer:

- **The number of requests.** Every turn is a request, and the client waits for all of them. `execution.maxIterations` caps it, [10 by default](https://frontmcp.dev/learn/giving-an-agent-tools#limiting-the-loop).
- **What each request repeats.** Instructions, tool descriptions and schemas are sent every turn. Give an agent the tools its task needs, not every tool you have.
- **What tools return.** Every result stays in the conversation until the end. A tool an agent uses should return what the model needs, not the whole record: the same advice as for [shaping tool results](https://frontmcp.dev/learn/shaping-tool-results), and it matters more here, because the result is paid for again on every later turn.
- **Where it runs.** An agent spends your key on behalf of whoever calls it. Put agents behind [authentication](https://frontmcp.dev/learn/authenticating-clients) and [rate limits](https://frontmcp.dev/learn/limiting-calls).

## The other direction: `connectOpenAI()` and friends

FrontMCP also exports `connectOpenAI()`, `connectClaude()`, `connectLangChain()` and `connectVercelAI()`. Despite the names, they don't connect an agent to a model. They go the other way: they connect *your* code to a FrontMCP server, in-process, and hand you its tools in the shape that provider's API or library expects. They're for when your own program runs the model and the loop, and the FrontMCP server is only the tools. [Running FrontMCP Anywhere](https://frontmcp.dev/learn/running-frontmcp-anywhere#calling-your-tools-through-a-client-connect) shows them.

| You want | Use |
| --- | --- |
| Your MCP server to run a model loop of its own, as a tool clients call | `@Agent` with `llm` |
| Your own program to run the model loop, with a FrontMCP server's tools | `connectClaude()`, `connectOpenAI()`, … |

## Recap

- `llm: { provider: "anthropic" | "openai", model, apiKey: { env } }` connects an agent to a real model. Install `@anthropic-ai/sdk` or `openai`; FrontMCP loads it on the first request.
- `apiKey: { env }` is read when the server starts, and a missing variable stops the server from starting. Keep keys out of your source.
- `AnthropicAdapter` and `OpenAIAdapter` are what `provider` builds. Build one yourself for `api: "responses"`, `baseUrl`, your own `client`, or `maxRetries`.
- Adapters retry a failed request 3 times, waiting 1, 2 and 4 seconds, including failures that can't succeed, like a wrong key. A call that still fails is a `TOOL_EXECUTION_ERROR`.
- Every turn is a request carrying the whole conversation. Wrap the adapter, or override the agent's `completion()`, to see `usage`. A reply cut off at `maxTokens` becomes the answer, silently.
- `connectOpenAI()` and friends connect your own code to a FrontMCP server; they aren't model adapters.

## Try some challenges

Each challenge runs hidden checks against your code. Edit the code, then press **Check**.

### Challenge: Fail fast when the provider is busy
This agent is called directly by clients, and the provider is rate-limiting every request. With `maxRetries: 2`, each call waits three seconds before it fails. Make a rate-limited call fail at once, after a single request.

```ts triage.agent.ts active
import { Agent, AgentContext, OpenAIAdapter, z } from "@frontmcp/sdk";
import { openai } from "./openai.example";

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  inputSchema: { ticketId: z.string() },
  llm: { adapter: new OpenAIAdapter({ model: "gpt-5", client: openai, maxRetries: 2 }) },
})
export class Triage extends AgentContext {}
```

```ts triage.agent.ts solution
import { Agent, AgentContext, OpenAIAdapter, z } from "@frontmcp/sdk";
import { openai } from "./openai.example";

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  inputSchema: { ticketId: z.string() },
  llm: { adapter: new OpenAIAdapter({ model: "gpt-5", client: openai, maxRetries: 0 }) },
})
export class Triage extends AgentContext {}
```

```ts openai.example.ts
// Stands in for the openai client, on a bad day: every request is rate-limited.

export let requests = 0;

export const openai = {
  chat: {
    completions: {
      async create(): Promise<any> {
        requests++;
        throw Object.assign(new Error("429 Rate limit reached for gpt-5"), { status: 429 });
      },
    },
  },
  responses: {
    async create(): Promise<any> {
      throw new Error("This example uses Chat Completions.");
    },
  },
};
```

```ts fast.test.ts hidden
import { test, expect } from "@frontmcp/testing";
import { requests } from "./openai.example";

test("the call fails with the provider's message", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result).toBeError("TOOL_EXECUTION_ERROR");
  expect(result.text()).toContain("429 Rate limit reached");
});

test("it fails within half a second", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result.durationMs).toBeLessThan(500);
});

test("the provider gets one request per call", async ({ mcp }) => {
  const before = requests;
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(requests - before).toBe(1);
});
```

**Hint:**
The waits come from the adapter, not from the agent's `execution` options. Count how many requests `maxRetries: 2` makes.

**Solution:**
`maxRetries` is how many times the adapter tries again after the first request, so `0` means one request and no waiting. The client's model gets the failure straight away and can tell the user, instead of holding the call for 1 + 2 seconds of retries that, during an outage, won't help. Jobs and other background work can afford more retries than a call someone is waiting on.

### Challenge: Report the tokens each call used
Support wants to see what triage costs. Add `tokens` to the agent's result: the total tokens, input and output, of every request that call made, even when two calls run at the same time. The model's answer should still be in `response`.

```ts triage.agent.ts active
import { Agent, AgentContext, AnthropicAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import { anthropic } from "./anthropic.example";

@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string() } })
class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
  inputSchema: { ticketId: z.string() },
  tools: [GetTicket],
  llm: { adapter: new AnthropicAdapter({ model: "claude-opus-5", client: anthropic }) },
})
export class Triage extends AgentContext {}
```

```ts triage.agent.ts solution
import { Agent, AgentContext, AnthropicAdapter, Tool, ToolContext, z } from "@frontmcp/sdk";
import type { AgentCompletion, AgentCompletionOptions, AgentPrompt, AgentToolDefinition } from "@frontmcp/sdk";
import { anthropic } from "./anthropic.example";

@Tool({ name: "get_ticket", description: "Get a support ticket by its id.", inputSchema: { id: z.string() } })
class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@Agent({
  name: "triage",
  description: "Triage a support ticket. Pass the ticket id.",
  systemInstructions: "You triage support tickets. Read the ticket with get_ticket, then give its priority and why.",
  inputSchema: { ticketId: z.string() },
  tools: [GetTicket],
  llm: { adapter: new AnthropicAdapter({ model: "claude-opus-5", client: anthropic }) },
})
export class Triage extends AgentContext {
  private tokens = 0;

  protected async completion(prompt: AgentPrompt, tools?: AgentToolDefinition[], options?: AgentCompletionOptions): Promise<AgentCompletion> {
    const reply = await super.completion(prompt, tools, options);
    this.tokens += reply.usage?.totalTokens ?? 0;
    return reply;
  }

  async execute(input: { ticketId: string }) {
    const answer = await super.execute(input);
    return { ...answer, tokens: this.tokens };
  }
}
```

```ts anthropic.example.ts
// Stands in for the Anthropic client: a tool call, then an answer, with the
// usage a real response reports.

export const anthropic = {
  messages: {
    async create(params: any): Promise<any> {
      if (params.messages.length === 1) {
        return {
          content: [{ type: "tool_use", id: "toolu_01", name: "get_ticket", input: { id: "T-1" } }],
          stop_reason: "tool_use",
          usage: { input_tokens: 600, output_tokens: 40 },
        };
      }
      return {
        content: [{ type: "text", text: "High: the customer can't log in." }],
        stop_reason: "end_turn",
        usage: { input_tokens: 700, output_tokens: 15 },
      };
    },
  },
};
```

```ts tokens.test.ts hidden
import { test, expect } from "@frontmcp/testing";

test("the result has `tokens`: 600 + 40 + 700 + 15", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result.json().tokens).toBe(1355);
});

test("it's counted per call, not in total", async ({ mcp }) => {
  await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result.json().tokens).toBe(1355);
});

test("two calls at once count their own", async ({ mcp }) => {
  const [a, b] = await Promise.all([mcp.tools.call("invoke_triage", { ticketId: "T-1" }), mcp.tools.call("invoke_triage", { ticketId: "T-1" })]);
  expect([a.json().tokens, b.json().tokens]).toEqual([1355, 1355]);
});

test("the answer is still in `response`", async ({ mcp }) => {
  const result = await mcp.tools.call("invoke_triage", { ticketId: "T-1" });
  expect(result.json().response).toBe("High: the customer can't log in.");
});
```

**Hint:**
The loop drops each reply's `usage`, so catch it on the way, in a method the loop calls for every request. Keep the count where each call has its own, and add it to the result where the agent's result is made.

**Solution:**
The loop calls the agent's `completion()` for every request, so an override sees every reply, and adds up `usage.totalTokens`, which `AnthropicAdapter` fills in from the response's `input_tokens` and `output_tokens`. FrontMCP creates the agent's class for each call, so `this.tokens` belongs to one call, even when two run at once. Overriding `execute()` adds it to the result after `super.execute(input)`. A wrapping adapter, as in [Cost and latency](#cost-and-latency), sees the same replies, but it's shared by every call, so it can't tell two calls at once apart.
