# Observability and telemetry

> Trace every request with OpenTelemetry, add your own spans and counters, serve /metrics, and answer health checks.

Source: https://frontmcp.dev/reference/server/observability

Three things let you watch a FrontMCP server in production. **Traces**: with `@frontmcp/observability` installed, `@FrontMcp({ observability })` records an OpenTelemetry span for each request, each call and each outbound `this.fetch()`, and `this.telemetry` adds your own. **Metrics**: `@FrontMcp({ metrics })` serves process gauges and your counters at `/metrics`, for Prometheus, and a meter provider you register sends the same counters to an OpenTelemetry collector. **Health checks**: `/healthz` and `/readyz` tell an orchestrator whether to restart the server or send it traffic. Tracing and metrics send data to systems outside the process, so most of this page runs in your project, not in the Playground; each section says what the Playground can show.

```ts
@FrontMcp({
  info, apps,
  observability: true,          // or { tracing, logging, requestLogs }
  metrics: { enabled: true },   // GET /metrics
  health: { probes },           // GET /healthz, /readyz
})
```

---

## Reference

### `observability`

`observability` needs the `@frontmcp/observability` package, an optional peer of `@frontmcp/sdk`, and it records spans through whatever OpenTelemetry tracer provider you register. FrontMCP doesn't register one for you.

```bash
npm install @frontmcp/observability @opentelemetry/api @opentelemetry/sdk-trace-base @opentelemetry/exporter-trace-otlp-http
```

```ts main.ts
import "reflect-metadata";
import { trace } from "@opentelemetry/api";
import { BasicTracerProvider, BatchSpanProcessor } from "@opentelemetry/sdk-trace-base";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { FrontMcp } from "@frontmcp/sdk";
import { HelpDeskApp } from "./help-desk.app";

// Any OpenTelemetry setup that registers a global tracer provider works. This one sends spans to a collector.
trace.setGlobalTracerProvider(
  new BasicTracerProvider({
    spanProcessors: [new BatchSpanProcessor(new OTLPTraceExporter({ url: "http://localhost:4318/v1/traces" }))],
  }),
);

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: true,
})
export default class Server {}
```

Register the provider before the server handles requests. Sending spans to a collector needs one running, so it can't be shown here; everything below about spans was checked by registering a provider that keeps them in memory.

`@frontmcp/observability` also exports `setupOTel({ serviceName?, exporter?, endpoint?, serviceVersion? })`, which registers a provider through OpenTelemetry's Node SDK and returns an async function that shuts it down. `exporter` is `"console"` (the default) or `"otlp"`, which sends to `endpoint` plus `/v1/traces`; `endpoint` defaults to `OTEL_EXPORTER_OTLP_ENDPOINT`, then `http://localhost:4318`, and `serviceName` to `OTEL_SERVICE_NAME`, then `"frontmcp-server"`. It needs `@opentelemetry/sdk-node`, `@opentelemetry/resources` and `@opentelemetry/semantic-conventions`, and `@opentelemetry/exporter-trace-otlp-http` for `"otlp"`; without them it throws `setupOTel: @opentelemetry/sdk-node is required`. Only those errors were checked here. The environment variables do nothing by themselves: they're read by `setupOTel()` and the `otlp` log sink, so setting them without calling `setupOTel()` exports nothing.

| Option | Type | Default | Description |
| --- | --- | --- | --- |
| `tracing` | `boolean \| TracingOptions` | `true` | Spans for requests and calls, and `this.telemetry`. `observability: true` is `{ tracing: true }`. |
| `logging` | `boolean \| { sinks, redactFields, includeStacks, staticFields }` | `false` | [Structured log entries](#structured-logs): each line of the [server's log](https://frontmcp.dev/reference/server/logging) as an object, sent to the sinks you list. |
| `requestLogs` | `boolean \| { maxEntries, includeSummaries, onRequestComplete }` | `false` | Gathers each HTTP request's log lines and outcome into one object, and passes it to `onRequestComplete`. See [request logs](#request-logs). |

Without the package, the server starts anyway and writes one warning: [`observability config is set but @frontmcp/observability is not installed`](#observability-config-is-set-but-frontmcpobservability-is-not-installed).

#### `TracingOptions`

| Option | Default | What it does |
| --- | --- | --- |
| `httpSpans` | `true` | A server span for each HTTP request, named after it: `POST /`. |
| `executionSpans` | `true` | A span for each call (`tools/call`, `resources/read`, `prompts/get`), and one inside it for the tool, resource or prompt that ran. |
| `fetchSpans` | `true` | A client span for each [`this.fetch()`](https://frontmcp.dev/reference/sdk/fetch), named after the method: `GET`. The `traceparent` that `this.fetch()` sends names this span; with `false`, it names the tool's span. |
| `transportSpans` | `true` | On FrontMCP's HTTP server, a `transport streamable-http` span for each message of a session opened by a client on a protocol version before 2026-07-28. |
| `authSpans` | `true` | Meant for authentication flows. A public or static-token server records none. |
| `flowStageEvents` | `true` | The `stage.*` events on the request and call spans, like `stage.findTool`. `false` leaves them out. |
| `hookSpans` | `false` | A `hook <stage>` span, like `hook willExecute`, for each hook a plugin, app or entry runs, inside its call's span. Its attributes are `frontmcp.hook.stage`, `frontmcp.hook.owner` (the plugin's class name) and `frontmcp.flow.name`. |
| `oauthSpans` | `true` | A span for each request to an OAuth endpoint, named after it: `oauth/token`. |
| `elicitationSpans` | `true` | Spans for elicitation requests and results. Not checked here. |
| `startupReport` | `true` | A `frontmcp.startup` span once the server is ready, with `frontmcp.startup.tools_count`, `resources_count`, `prompts_count`, `plugins_count` and `duration_ms`. |

Changed in 1.9.3: before, `flowStageEvents`, `hookSpans`, `oauthSpans`, `elicitationSpans` and `startupReport` were accepted but not read.

#### Spans

A 2026-07-28 `tools/call` records these spans. Each span's trace id is the request's [`this.context.traceContext.traceId`](https://frontmcp.dev/reference/sdk/context), so a client that sends `traceparent` sees them in its own trace.

| Span | Kind | Parent | Attributes | Events |
| --- | --- | --- | --- | --- |
| `POST /` | server | The caller's span, from `traceparent` | `http.request.method`, `url.path`, `http.response.status_code`, `frontmcp.request.id`, `frontmcp.scope.id`, `mcp.session.id` | `stage.traceRequest`, `stage.checkAuthorization`, `stage.router`, `stage.finalize`, … |
| `tools/call` | server | `POST /` | `rpc.system` (`"mcp"`), `rpc.method`, `mcp.method.name`, `mcp.session.id`, `auth.authorities.result` | `stage.parseInput`, `stage.findTool`, `stage.validateInput`, `stage.validateOutput`, `stage.finalize`, … |
| `tool close_ticket` | internal | `tools/call` | `frontmcp.tool.name`, `mcp.component.type` (`"tool"`), `mcp.component.key` (`"tool:close_ticket"`), `enduser.id`, `enduser.scope` | `stage.execute.start`, `stage.execute.done`, and yours |
| `GET` | client | `tools/call` | `http.request.method`, `url.full`, `http.response.status_code` | |

A resource read gives `resources/read` and `resource tickets://T-1`; a prompt, `prompts/get` and `prompt triage`; a `tools/list`, one `tools/list` span. Each stage event has a `duration_ms` attribute. `mcp.session.id` is a hash of the session id, never the id itself. The `traceparent` that `this.fetch()` sends names its `GET` span, so the service you call records its work under it. (Changed in 1.9.3: before, `tools/call` was a sibling of `POST /`, and `this.fetch()` passed on the caller's `traceparent` unchanged.)

#### Caveats

- **A failed call ends its spans with status ERROR**, with what the client would get in production, whether the error was thrown or passed to `this.fail()`. The `tool …` span and the `tools/call` span both get the public message as their status message, and an `exception` event whose type is the error's code: `no such ticket` and `PUBLIC_ERROR` for a `PublicMcpError`, and for any other error, FrontMCP's `Internal FrontMCP error. Please contact support with error ID: err_…`, with the error id the client got, and `TOOL_EXECUTION_ERROR` when the tool threw it or `SERVER_ERROR` when it was passed to `this.fail()`. The event's stack still has the error's own message. See [A failed span says `Internal FrontMCP error`](#a-failed-span-says-internal-frontmcp-error). (Changed in 1.9.3: before, the status message was the error's own message, and empty after `this.fail()`.)
- **ES module projects record the same.** Changed in 1.9.4: in a project with `"type": "module"`, every failure was recorded as `SERVER_ERROR`, a `PublicMcpError` too, with an error id the client didn't get: see [the troubleshooting entry](#a-failed-span-says-internal-frontmcp-error).
- When no tracer provider is registered as the server starts, and `NODE_ENV` isn't `development`, FrontMCP logs [`observability: tracing enabled but no TracerProvider configured`](#observability-tracing-enabled-but-no-tracerprovider-configured). In development it registers a provider that prints each span to the console instead, one line per span: `↳ SPAN tools/call [04bac610] 2.5ms SERVER rpc.system=mcp rpc.method=tools/call …`. (Changed in 1.9.3: before, that failed silently on `@opentelemetry/sdk-trace-base` 2.x.)

### `this.telemetry`

With `observability` on, tools, resources and prompts have `this.telemetry`, for spans and events of your own. What you add lands in the same trace, under the span of the tool, resource or prompt that's running.

```ts search.tool.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";
import { TicketIndex } from "./ticket-index";

@Tool({ name: "search_tickets", description: "Search tickets by words in their title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    this.telemetry.addEvent("query-parsed", { terms: query.split(" ").length });
    const tickets = await this.telemetry.withSpan("ticket-index.search", async (span) => {
      span.setAttribute("db.query", query);
      return this.get(TicketIndex).search(query);
    });
    this.telemetry.setAttributes({ "results.count": tickets.length });
    return { tickets };
  }
}
```

The call records a `ticket-index.search` span, with `db.query`, as a child of `tool search_tickets`, and the tool's span gets the `query-parsed` event and the `results.count` attribute.

| Member | Description |
| --- | --- |
| `withSpan(name, fn, attributes?)` | Runs `fn(span)` inside a new child span and ends it. If `fn` throws, the span ends with status ERROR and an `exception` event, and the error is rethrown. |
| `startSpan(name, attributes?)` | Starts a child span and returns a `TelemetrySpan`. Call `end()` or `endWithError()` yourself: a span that isn't ended is never exported. |
| `addEvent(name, attributes?)` | Adds an event to the running tool's, resource's or prompt's span. |
| `setAttributes(attributes)` | Sets attributes on that span. |
| `createCounter(name, description?)` | A [counter](#counting-things) for `/metrics` and for [an OpenTelemetry collector](#sending-counters-to-an-opentelemetry-collector). |
| `traceId` | The request's trace id, the same as `this.context.traceContext.traceId`. |
| `sessionId` | A 16-character hash of the session id, the value of `mcp.session.id`. |

Spans you start have `frontmcp.request.id`, `frontmcp.scope.id` and `mcp.session.id` besides your own attributes. Attribute values are strings, numbers or booleans.

| `TelemetrySpan` member | Description |
| --- | --- |
| `setAttribute(key, value)`, `setAttributes(attributes)` | Set attributes. Return the span, for chaining. |
| `addEvent(name, attributes?)` | Add an event. |
| `recordError(error)` | Record an `exception` event and set status ERROR. A later `end()` keeps the ERROR. |
| `end()` | End the span, with status OK unless an error was recorded. |
| `endWithError(error)` | End the span with status ERROR. |
| `raw` | The OpenTelemetry `Span`, for anything else. |

`this.telemetry` exists only while tracing is on. Its type comes from `@frontmcp/observability`: TypeScript knows it once a file in your project imports the package, for `createCounter()` or with a bare `import "@frontmcp/observability";`. That includes prompts: `PromptContext` has it too.

It needs `@frontmcp/observability`, which the Playground can't load, so the examples of it on this page are code to run in your project.

#### Caveats

- With `tracing: false`, reading `this.telemetry` throws [`ObservabilityPlugin is not installed or tracing is disabled`](#observabilityplugin-is-not-installed-or-tracing-is-disabled).
- Without the package, `this.telemetry` is `undefined`.
- Without a registered tracer provider, every span is a no-op: the code runs, and nothing is recorded.

### Structured logs

`observability.logging` turns each line of the server's log into an object with the request's ids, and sends it to sinks. It adds to the [console and your transports](https://frontmcp.dev/reference/server/logging); it doesn't replace them. Like the rest of `observability`, it doesn't run in the Playground; for a format of your own that does, write a [log transport](https://frontmcp.dev/reference/server/logging#writing-a-transport).

```ts main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: {
    logging: {
      sinks: [{ type: "stdout" }],
      redactFields: ["password", "token", "authorization"],
      staticFields: { service: "help-desk", region: "eu-west-1" },
    },
  },
  logging: { enableConsole: false },
})
export default class Server {}
```

```json
{"service":"help-desk","region":"eu-west-1","timestamp":"2026-10-06T13:52:55.409Z","level":"info","severity_number":9,"message":"closing T-1","prefix":"tool:close_ticket","trace_id":"0af7651916cd43dd8448eb211c80319c","span_id":"b7ad6b7169203331","trace_flags":1,"request_id":"8f691bc7-290c-4f3f-ace9-0deadc6fc4fe","session_id_hash":"bef6cb302a57","scope_id":"root","flow_name":"tools:call-tool","elapsed_ms":10}
```

| Option | Default | Description |
| --- | --- | --- |
| `sinks` | none | Where entries go. **Required**: `logging: true` alone has no sinks and writes nothing. |
| `redactFields` | `[]` | Keys to replace with `"[REDACTED]"` in the objects passed to the logger, in nested objects too, ignoring case. |
| `includeStacks` | `true` | Include the stack of an `Error` passed to the logger. |
| `staticFields` | `{}` | Fields added to every entry. |

| Sink | Writes |
| --- | --- |
| `{ type: "stdout" }` | One JSON object per line to stdout, or to stderr for a stdio server. |
| `{ type: "console" }` | A readable line: `[10:41:05.879] ERROR [716c9764:a556272f] [tool:pay] charge failed { orderId=O-1 } [PublicMcpError: card declined] (6ms)`, with the trace and request ids. |
| `{ type: "callback", fn }` | Calls `fn(entry)`. |
| `{ type: "otlp", endpoint?, headers?, batchSize?, flushIntervalMs?, serviceName? }` | OTLP over HTTP to a collector; `endpoint` defaults to `OTEL_EXPORTER_OTLP_ENDPOINT`, then `http://localhost:4318`. Not checked here. |
| `{ type: "winston", logger }`, `{ type: "pino", logger }` | Your winston or pino logger. Not checked here. |

Each entry has `timestamp`, `level` (lowercase), `severity_number` (OpenTelemetry's: 5, 6, 9, 13 or 17), `message` and `prefix`. Lines written during a request add `trace_id`, `span_id` (the caller's span), `trace_flags`, `request_id`, `session_id_hash` and `scope_id`, `flow_name`, the flow that wrote the line (`http:request`, `session:verify`, `tools:call-tool`, …), and `elapsed_ms` since the request started. Objects passed after the message are merged into `attributes`; an `Error` becomes `error: { type, message, error_id, stack }`.

### Request logs

`observability: { requestLogs }` gathers what happened during each HTTP request into one object, and hands it to `onRequestComplete` once the response is ready: one per request, failed ones included. It needs `@frontmcp/observability`, and works whether or not `observability.logging` has sinks.

```ts main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: {
    requestLogs: {
      onRequestComplete: (log) => {
        if (log.status === "error") alerts.send(log); // your own code
      },
    },
  },
})
export default class Server {}
```

| Option | Default | Description |
| --- | --- | --- |
| `onRequestComplete` | | Called with the request's log. A promise it returns is awaited, and an error it throws is ignored. |
| `maxEntries` | `500` | Log lines kept per request. Later ones are dropped. |
| `includeSummaries` | | Accepted, but not read by FrontMCP 1.9.4. |

| Field | What it holds |
| --- | --- |
| `request_id`, `trace_id`, `session_id_hash`, `scope_id` | The request's ids, as on [structured log](#structured-logs) entries. |
| `start_time`, `end_time`, `duration_ms` | When it ran. |
| `http_method`, `http_path`, `rpc_method` | `"POST"`, `"/"`, `"tools/call"`. |
| `tool_name`, `resource_uri`, `prompt_name` | What it called, read or got. |
| `status`, `status_code` | `"ok"` or `"error"`, and the HTTP status the server answered with: `200` for a tool call that failed too, `401` for a caller it refused. (Changed in 1.9.3: before, `status_code` was only set from 400.) |
| `error` | For a failed call, `{ type, message, code, error_id }`, as the client gets it in production, whether the error was thrown or passed to `this.fail()`: `{ type: "PublicMcpError", message: "no such ticket", code: "PUBLIC_ERROR", error_id: "err_…" }` for a public error, and for any other, `{ type: "ToolExecutionError", message: "Internal FrontMCP error. Please contact support with error ID: err_…", code: "TOOL_EXECUTION_ERROR", error_id: "err_…" }` when a tool threw it, or `GenericServerError` and `SERVER_ERROR` when it was passed to `this.fail()`. `error_id` is the `errorId` the client got. (Changed in 1.9.3: before, after `this.fail()` it was `{ type: "Error", message: "" }`, and another error's own message was recorded.) |
| `entries` | Each line written during the request: `{ timestamp, level, message, stage, elapsed_ms, attributes? }`, where `stage` is the flow, as `flow_name` above. |
| `authenticated`, `auth_type` | `true` and `"bearer"` for a caller the server verified with a token; `false` and `"anonymous"` for an anonymous caller, or one it refused. (Changed in 1.9.3: before, `authenticated` was always `false` and `auth_type` missing.) |
| `hooks_triggered` | Each hook a plugin, app or entry ran during the request, as `<flow>:<stage>`: `["tools:call-tool:willExecute"]`. (Changed in 1.9.3: before, it was always `[]`.) |

This was checked in Node, with `createFetchHandler()`; the Playground can't load `@frontmcp/observability`.

### `metrics`

`metrics` adds a `/metrics` endpoint to FrontMCP's own HTTP server, and, since 1.9.0, to [`createFetchHandler()`](https://frontmcp.dev/reference/sdk/create-fetch-handler). It's off by default, and needs `@frontmcp/observability` with `@opentelemetry/sdk-trace-base`: without them, the server doesn't start (`Cannot find module '@frontmcp/observability'`, or `'@opentelemetry/sdk-trace-base'`).

```ts main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  metrics: { enabled: true, auth: "token" }, // reads FRONTMCP_METRICS_TOKEN
})
export default class Server {}
```

| Option | Type | Default | Description |
| --- | --- | --- | --- |
| `enabled` | `boolean` | `false` | Serve the endpoint. |
| `path` | `string` | `"/metrics"` | Where. It can't be `/mcp`, `/sse` or `/messages`, or start with one of them: the server doesn't start, with [`MetricsPathConflictError`](#metricspath--conflicts-with-an-mcp-transport-route). |
| `format` | `"prometheus" \| "json"` | `"prometheus"` | Prometheus text, or `{ counters, gauges }` as JSON. |
| `auth` | `"public" \| "token" \| { token }` | `"public"` | `"token"` requires `Authorization: Bearer <token>`, with the token read from `tokenEnv` at startup. `{ token: "…" }` puts the token in the code, which is fine for local testing only. |
| `tokenEnv` | `string` | `"FRONTMCP_METRICS_TOKEN"` | The environment variable that holds the token. If it isn't set, the server doesn't start: [`MetricsTokenNotConfiguredError`](#metricsauth-is-token-but-env-var--is-not-set). |
| `include` | `MetricsCategory[]` | every category | Only metrics whose names start with the categories' prefixes: `process` (`frontmcp_process_`, `frontmcp_nodejs_`), `tools` (`frontmcp_tool_`), `resources` (`frontmcp_resource_`), `http` (`frontmcp_http_`), `storage` (`frontmcp_storage_`), `skills` (`frontmcp_skills_`), `auth` (`frontmcp_auth_`), `sessions` (`frontmcp_session_`). |
| `process` | `{ eventLoopLag?, fdCount?, activeHandles? }` | all `true` | Which Node gauges to collect. `fdCount` is Linux only. |

| Request | Response |
| --- | --- |
| `GET /metrics` | `200`, `Content-Type: text/plain; charset=utf-8; version=0.0.4` (or `application/json`), `Cache-Control: no-store`. |
| Without `Authorization`, with `auth: "token"` | `401` `{ "error": "unauthorized", "message": "Missing or malformed Authorization header" }` |
| With the wrong token | `403` `{ "error": "forbidden", "message": "Bearer token did not match the configured metrics token" }` |

The scrape holds the process gauges and [your counters](#counting-things):

```text
# HELP frontmcp_process_resident_memory_bytes Resident memory size in bytes
# TYPE frontmcp_process_resident_memory_bytes gauge
frontmcp_process_resident_memory_bytes 158318592
# HELP frontmcp_process_cpu_seconds_total CPU time consumed since collector start, by mode (seconds)
# TYPE frontmcp_process_cpu_seconds_total gauge
frontmcp_process_cpu_seconds_total{mode="user"} 0.022128
# TYPE frontmcp_nodejs_eventloop_lag_seconds gauge
frontmcp_nodejs_eventloop_lag_seconds{quantile="p99"} 5.11e-7
# TYPE helpdesk_tickets_closed_total counter
helpdesk_tickets_closed_total{priority="high"} 2
```

The gauges are `frontmcp_process_resident_memory_bytes`, `_heap_bytes`, `_heap_used_bytes`, `_external_bytes`, `_uptime_seconds` and `_cpu_seconds_total{mode}`, and `frontmcp_nodejs_eventloop_lag_seconds{quantile="p99"}`, `_active_handles`, `_active_requests` and, on Linux, `_open_fds`.

#### Skilled OpenAPI's counters

[Skilled OpenAPI](https://frontmcp.dev/reference/plugins/skilled-openapi) counts its bundles: `frontmcp_skills_bundle_pulls_total` (`status`, `source`) each time it applies one, and, when it checks a bundle's signature, `frontmcp_skills_signature_verifications_total` (`status`) and `frontmcp_skills_signature_failures_total` (`reason`, like `missing_integrity`). Each bundle it applies also records a `skill.bundle.swap` span, with `bundle_id`, `version` and `skill_count`. It does so only when `ObservabilityPlugin` from `@frontmcp/observability` comes before it in `plugins`:

```ts main.ts
import { ObservabilityPlugin } from "@frontmcp/observability";
import { SkilledOpenApiPlugin } from "@frontmcp/plugin-skilled-openapi";

@FrontMcp({
  info: { name: "billing", version: "1.0.0" },
  apps: [],
  metrics: { enabled: true },
  plugins: [ObservabilityPlugin.init({ tracing: true }), SkilledOpenApiPlugin.init({ source, trustedKeys })],
})
export default class Server {}
```

```text
# TYPE frontmcp_skills_bundle_pulls_total counter
frontmcp_skills_bundle_pulls_total{source="unknown",status="ok"} 1
```

With `observability: true` instead, or with `ObservabilityPlugin` after it, the plugin can't reach the counters and records nothing, though tracing works. This was checked in Node with `createFetchHandler()`, an `inline` bundle and an in-memory span exporter.

#### Caveats

- **FrontMCP 1.9.4 counts no requests, tool calls or errors.** The scrape is the process gauges and the counters you create, plus [Skilled OpenAPI](https://frontmcp.dev/reference/plugins/skilled-openapi)'s own when it's installed, and only then: see [below](#skilled-openapis-counters). The `include` categories other than `process` and `skills` only match counters you name with their prefixes.
- **Every server FrontMCP builds serves it**: the one `@FrontMcp` and `FrontMcpInstance.bootstrap()` start, the handler `FrontMcpInstance.createHandler()` returns for serverless platforms, and [`createFetchHandler()`](https://frontmcp.dev/reference/sdk/create-fetch-handler)'s, which before 1.9.0 ignored `metrics` and answered `/metrics` with `404`. The Playground can't load `@frontmcp/observability`, so it can't show it.
- A counter's description isn't printed as a `# HELP` line.

### `health`

FrontMCP answers health checks without any configuration. `health` adds your own checks, or changes the paths. The server's other options are on [`@FrontMcp`](https://frontmcp.dev/reference/sdk/frontmcp#options).

```ts main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  health: {
    probes: [
      {
        name: "tickets-db",
        async check() {
          const started = Date.now();
          await db.query("select 1");
          return { status: "healthy", latencyMs: Date.now() - started };
        },
      },
    ],
    readyz: { timeoutMs: 2000 },
  },
})
export default class Server {}
```

| Option | Type | Default | Description |
| --- | --- | --- | --- |
| `enabled` | `boolean` | `true` | `false` removes `/healthz`, `/readyz` and `/health`. |
| `healthzPath` | `string` | `"/healthz"` | The liveness path. `/health` answers the same whatever this is. |
| `readyzPath` | `string` | `"/readyz"` | The readiness path. |
| `probes` | `{ name, check() }[]` | `[]` | Your checks, run in parallel on every `/readyz`. `check()` returns `{ status, latencyMs?, details?, error? }`, with `status` `"healthy"`, `"degraded"` or `"unhealthy"`. |
| `includeDetails` | `boolean` | `true` in development, `false` in production | Include each probe's result in `/readyz`. |
| `readyz` | `{ enabled?, timeoutMs? }` | on, unless the runtime is an edge one or the deployment is serverless; `timeoutMs: 5000` | `enabled: false` removes `/readyz`: it's a `404`. A probe that takes longer than `timeoutMs` is `unhealthy`, with `"Probe timed out after 2000ms"`. |

| Request | FrontMCP's HTTP server (and `createHandler()`) answers |
| --- | --- |
| `GET /healthz`, `GET /health` | `200` `{ "status": "ok", "server": { "name", "version" }, "runtime": { "platform", "runtime", "deployment", "env" }, "uptime": 3600.5 }`. No checks: if the process can answer, it's alive. |
| `GET /readyz` | `200` `{ "status": "ready", "totalLatencyMs", "catalog": { "toolsHash", "toolCount", "resourceCount", "promptCount", "skillCount", "agentCount" }, "probes": { … } }`, or `503` with `"status": "not_ready"` when a probe is `unhealthy`, throws (its message becomes `error`) or times out. `degraded` still counts as ready. |

The `probes` it lists include FrontMCP's own `session-store` check next to yours, and a `remote:<app id>` check for each [remote app](https://frontmcp.dev/reference/server/remote), which is `degraded`, and so still ready, until the first check of the remote has run. `toolsHash` is a hash of the tool names, the same on every copy of a server with the same tools, so a replica that's out of step stands out.

[`createFetchHandler()`](https://frontmcp.dev/reference/sdk/create-fetch-handler) honours `health` too: `enabled`, both paths, `probes`, `includeDetails` and `readyz`. Its `/readyz` answers as above, with `"transport": "web-fetch"` added, and `503` when a probe is `unhealthy`. Its `/healthz`, and `/health`, answer the shorter `200` `{ "status": "ok", "server": { "name", "version" }, "transport": "web-fetch" }`, with no `runtime` or `uptime`. That's what the Playground runs, so [the example below](#answering-health-checks) shows it. (Before 1.8.7 it ignored `health`: `/healthz` and `/readyz` both returned a fixed `200`, no probe ran, and `/health` was a `404`.)

### Trace context helpers

`@frontmcp/sdk` exports the functions it uses to read [W3C trace context](https://www.w3.org/TR/trace-context/). They work without `@frontmcp/observability`. Each returns `{ traceId, parentId, traceFlags, raw }`, where `raw` is a `traceparent` value.

| Function | Returns |
| --- | --- |
| `parseTraceContext(headers)` | The trace context of a request's headers: from `traceparent`, else from `x-frontmcp-trace-id` (32 hex characters) with a new parent id, else a new trace. A malformed or all-zero `traceparent` counts as missing. |
| `generateTraceContext()` | A new trace, sampled. |
| `createChildSpanContext(parent)` | The same trace with a new parent id: the `traceparent` to send for work you do on the trace's behalf. |

---

## Usage

### Following a request across services

A request's trace id comes from the caller. A `traceparent` header continues the caller's trace; without one, FrontMCP also accepts an `x-frontmcp-trace-id` header, and without either, the request starts a trace of its own. The same id is in [`this.context.traceContext`](https://frontmcp.dev/reference/sdk/context#joining-a-trace), on every line of [`this.contextLogger`](https://frontmcp.dev/reference/server/logging#telling-requests-apart) (its first eight characters), in the `traceparent` that [`this.fetch()`](https://frontmcp.dev/reference/sdk/fetch) sends on, and, with tracing on, in every span. That's what lets you go from a log line to the trace, and from the trace to the services it called. This part runs in the Playground:

```ts trace.tool.ts active
import { App, Tool, ToolContext } from "@frontmcp/sdk";

@Tool({ name: "trace_info", description: "Show this request's trace. For debugging.", inputSchema: {} })
export class TraceInfo extends ToolContext {
  async execute() {
    this.contextLogger.info("looking up the trace");
    const { traceId, parentId, traceFlags } = this.context.traceContext;
    return { traceId, parentId, traceFlags };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [TraceInfo] })
export class HelpDesk {}
```

```ts trace.test.ts
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance, parseTraceContext } from "@frontmcp/sdk";
import { HelpDesk } from "./trace.tool";

const traceparent = "00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01";

async function callWith(headers: Record<string, string>) {
  const handler = await FrontMcpInstance.createFetchHandler({ info: { name: "help-desk", version: "1.0.0" }, apps: [HelpDesk] });
  const response = await handler(
    new Request("https://desk.example.com/", {
      method: "POST",
      headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/call", "mcp-name": "trace_info", ...headers },
      body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/call", params: { name: "trace_info", arguments: {}, _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } } }),
    }),
  );
  return (await response.json()).result.structuredContent;
}

test("a traceparent header continues the caller's trace", async () => {
  expect(await callWith({ traceparent })).toEqual({ traceId: "0af7651916cd43dd8448eb211c80319c", parentId: "b7ad6b7169203331", traceFlags: 1 });
});

test("without one, x-frontmcp-trace-id sets the trace id", async () => {
  const trace = await callWith({ "x-frontmcp-trace-id": "4BF92F3577B34DA6A3CE929D0E0E4736" });
  expect(trace.traceId).toBe("4bf92f3577b34da6a3ce929d0e0e4736");
  expect(trace.parentId).toMatch(/^[0-9a-f]{16}$/);
});

test("a malformed traceparent starts a new trace", async () => {
  const trace = await callWith({ traceparent: "00-00000000000000000000000000000000-b7ad6b7169203331-01" });
  expect(trace.traceId).toMatch(/^[0-9a-f]{32}$/);
  expect(trace.traceId).not.toBe("00000000000000000000000000000000");
});

test("parseTraceContext() reads headers the same way", () => {
  expect(parseTraceContext({ traceparent })).toEqual({ traceId: "0af7651916cd43dd8448eb211c80319c", parentId: "b7ad6b7169203331", traceFlags: 1, raw: traceparent });
  expect(parseTraceContext({ "x-frontmcp-trace-id": "4bf92f3577b34da6a3ce929d0e0e4736" }).traceId).toBe("4bf92f3577b34da6a3ce929d0e0e4736");
});
```

The **Logs** tab shows the tagged line: its second id is the start of the `traceId` in the result. A `traceparent` inside the request's `_meta` isn't read; only the header is (see [Joining a trace](https://frontmcp.dev/reference/sdk/context#joining-a-trace)). Work outside a request, like a job, has no trace of its own, and `this.fetch()` sends no `traceparent` there: start one with `generateTraceContext()` and send its `raw` yourself.

### Adding your own spans

The spans FrontMCP records show which call was slow. Your own spans show which part of it: a database query, a call to a model, a loop over pages of results. Wrap each in `this.telemetry.withSpan()`, and record what a later reader will want to filter by as attributes:

```ts close-ticket.tool.ts
import { PublicMcpError, Tool, ToolContext, z } from "@frontmcp/sdk";
import { Tickets } from "./tickets";

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string() } })
export class CloseTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    const ticket = await this.telemetry.withSpan("tickets.load", () => this.get(Tickets).load(id), { "ticket.id": id });
    if (!ticket) this.fail(new PublicMcpError(`There's no ticket ${id}.`));

    this.telemetry.setAttributes({ "ticket.priority": ticket.priority });
    await this.telemetry.withSpan("tickets.close", async (span) => {
      await this.get(Tickets).close(id);
      span.addEvent("customer-notified");
    });
    return { id, status: "closed" };
  }
}
```

A call gives `tool close_ticket`, with the `ticket.priority` attribute, and inside it `tickets.load` and `tickets.close`, each timed. If `load()` throws, `tickets.load` ends with status ERROR and the error as an `exception` event, and the call fails as usual. This needs `@frontmcp/observability`, so it doesn't run in the Playground.

### Checking spans in a test

`@frontmcp/observability` has helpers for tests: `createTestTracer()` gives a tracer provider that keeps spans in memory, and `assertSpanExists()`, `assertSpanAttribute()`, `findSpan()` and `findSpansByAttribute()` search them. Register the provider globally, or FrontMCP's spans don't reach it. The server must run in the test's process, as it does with `createFetchHandler()`:

```ts close-ticket.spans.test.ts
import { trace } from "@opentelemetry/api";
import { assertSpanAttribute, assertSpanExists, createTestTracer } from "@frontmcp/observability";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDeskApp } from "./help-desk.app";

const { provider, exporter, cleanup } = createTestTracer();
trace.setGlobalTracerProvider(provider); // without this, the exporter stays empty

afterEach(() => exporter.reset());
afterAll(() => cleanup());

test("closing a ticket records a tickets.close span", async () => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDeskApp],
    observability: true,
  });
  await handler(closeTicketRequest("T-1")); // a 2026-07-28 tools/call request

  const spans = exporter.getFinishedSpans();
  assertSpanExists(spans, "tickets.close");
  assertSpanAttribute(assertSpanExists(spans, "tool close_ticket"), "mcp.component.key", "tool:close_ticket");
});
```

`assertSpanExists()` returns the span, or throws `Expected span "tickets.close" but found: [tool close_ticket, tools/call, POST /]`. This is a test for your own project's runner; the Playground's tests can't import `@frontmcp/observability`.

### Counting things

A counter adds up across requests, which spans can't: tickets closed, cache hits, calls to a paid API. Create it once, with a name in snake case ending in `_total`, and add to it with `inc(by?, attributes?)`. `/metrics` shows every counter, one line per set of attributes:

```ts close-ticket.tool.ts
import { createCounter } from "@frontmcp/observability";
import { Tool, ToolContext, z } from "@frontmcp/sdk";

const ticketsClosed = createCounter("helpdesk_tickets_closed_total", "Tickets closed");

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string(), priority: z.enum(["low", "high"]) } })
export class CloseTicket extends ToolContext {
  async execute({ id, priority }: { id: string; priority: "low" | "high" }) {
    ticketsClosed.inc(1, { priority });
    return { id, status: "closed" };
  }
}
```

```text
# TYPE helpdesk_tickets_closed_total counter
helpdesk_tickets_closed_total{priority="high"} 2
```

`this.telemetry.createCounter(name)` returns the same counter. Keep attribute values to a small, fixed set, like a priority or a status: each new value is a new series to store. Counts live in the process's memory, so each replica reports its own, and they start at zero on restart. `getMetricSnapshot()` and `getCounterTotal(name)` from `@frontmcp/observability` read them, which is handy in tests. This needs the package, so it doesn't run in the Playground.

### Sending counters to an OpenTelemetry collector

`/metrics` waits for Prometheus to scrape it. To push the same counters to an OpenTelemetry collector, register a meter provider: every counter also reports to OpenTelemetry's global meter, as the instrumentation scope `@frontmcp/observability`. Neither `observability` nor `metrics` registers one, since `observability` sets up tracing only, and neither option is needed: the counters come from `@frontmcp/observability` alone. Register the provider in a file of its own, and import it first:

```bash
npm install @opentelemetry/api @opentelemetry/sdk-metrics @opentelemetry/exporter-metrics-otlp-http @opentelemetry/resources
```

```ts telemetry.ts
import { metrics } from "@opentelemetry/api";
import { OTLPMetricExporter } from "@opentelemetry/exporter-metrics-otlp-http";
import { resourceFromAttributes } from "@opentelemetry/resources";
import { MeterProvider, PeriodicExportingMetricReader } from "@opentelemetry/sdk-metrics";

export const meterProvider = new MeterProvider({
  resource: resourceFromAttributes({ "service.name": "help-desk" }),
  readers: [
    new PeriodicExportingMetricReader({
      exporter: new OTLPMetricExporter({ url: "http://localhost:4318/v1/metrics" }),
      exportIntervalMillis: 10_000,
    }),
  ],
});

metrics.setGlobalMeterProvider(meterProvider);
```

```ts main.ts
import "reflect-metadata";
import "./telemetry"; // before any file that creates a counter
import { FrontMcp } from "@frontmcp/sdk";
import { HelpDeskApp } from "./help-desk.app";

@FrontMcp({ info: { name: "help-desk", version: "1.0.0" }, apps: [HelpDeskApp] })
export default class Server {}
```

With the `close_ticket` tool from [Counting things](#counting-things) in `HelpDeskApp`, two calls with `priority: "high"` reached the collector as `helpdesk_tickets_closed_total`, a sum of `2` with the attribute `priority: "high"`, from a resource whose `service.name` is `help-desk`. `meterProvider.forceFlush()` sends what has been counted so far, which a test can use instead of waiting for the interval.

The import order matters. `createCounter()` asks OpenTelemetry for the counter the first time it sees a name, and keeps it. A counter created before the provider is registered reports to a meter that does nothing, for as long as the process runs, while `/metrics` and `getCounterTotal()` still count it. A counter at the top of a tool's file is created when the file is imported, which is why `telemetry.ts` comes first. `this.telemetry.createCounter(name)` runs later, inside a call, but returns the counter the name already has.

`setupOTel()`, from the [`observability`](#observability) section, registers a meter provider too, through OpenTelemetry's Node SDK. With that SDK's defaults (checked with `@opentelemetry/sdk-node` 0.219), it sends the counters every 60 seconds, as OTLP over protobuf, to `OTEL_EXPORTER_OTLP_ENDPOINT` plus `/v1/metrics`, which is `http://localhost:4318/v1/metrics` when the variable isn't set. Its `endpoint` option moves the traces only. `OTEL_METRICS_EXPORTER=none` turns the meter provider off. Nothing reports that the collector isn't there.

This needs the package, so it doesn't run in the Playground. It was checked in Node with FrontMCP 1.9.4, with an in-memory exporter and with a small HTTP server standing in for the collector.

### Answering health checks

An orchestrator asks two questions. `/healthz`: is the process alive, or should it be restarted? `/readyz`: can it take traffic now? Point liveness at the first and readiness at the second, and put what the server can't work without, like its database, in `health.probes`. A failing probe turns `/readyz` into a `503`, and the orchestrator stops sending traffic until it passes, while `/healthz` stays `200`.

A server behind [`createFetchHandler()`](https://frontmcp.dev/reference/sdk/create-fetch-handler), like this Playground's, does the same:

```ts health.test.ts active
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";

const unhealthy = { name: "tickets-db", check: async () => ({ status: "unhealthy" as const, error: "connection refused" }) };

const requestTo = async (path: string, health: object = { probes: [unhealthy] }) => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDesk],
    health: { includeDetails: true, ...health },
  });
  return handler(new Request(`https://desk.example.com${path}`));
};

test("/healthz is 200 whatever the probes say", async () => {
  const response = await requestTo("/healthz");
  expect(response.status).toBe(200);
  expect(await response.json()).toEqual({ status: "ok", server: { name: "help-desk", version: "1.0.0" }, transport: "web-fetch" });
});

test("/readyz runs `health.probes` and answers 503 while one is unhealthy", async () => {
  const response = await requestTo("/readyz");
  expect(response.status).toBe(503);
  const body = await response.json();
  expect(body.status).toBe("not_ready");
  expect(body.probes["tickets-db"]).toEqual({ status: "unhealthy", error: "connection refused" });
  expect(body.probes["session-store"].status).toBe("healthy");
});

test("/readyz is 200 once the probes pass", async () => {
  const healthy = { name: "tickets-db", check: async () => ({ status: "healthy" as const, latencyMs: 3 }) };
  const response = await requestTo("/readyz", { probes: [healthy] });
  expect(response.status).toBe(200);
  expect((await response.json()).status).toBe("ready");
});

test("a probe that takes longer than `readyz.timeoutMs` is unhealthy", async () => {
  const slow = { name: "slow", check: () => new Promise<{ status: "healthy" }>((resolve) => setTimeout(() => resolve({ status: "healthy" }), 300)) };
  const response = await requestTo("/readyz", { probes: [slow], readyz: { timeoutMs: 50 } });
  expect(response.status).toBe(503);
  expect((await response.json()).probes.slow.error).toBe("Probe timed out after 50ms");
});

test("`healthzPath` and `readyzPath` move the paths, and `enabled: false` removes them", async () => {
  const moved = { healthzPath: "/live", readyzPath: "/ready", probes: [] };
  expect((await requestTo("/live", moved)).status).toBe(200);
  expect((await requestTo("/ready", moved)).status).toBe(200);
  expect((await requestTo("/healthz", moved)).status).toBe(404);
  expect((await requestTo("/healthz", { enabled: false })).status).toBe(404);
});
```

```ts help-desk.app.ts
import { App, Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string() } })
class CloseTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, status: "closed" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [CloseTicket] })
export class HelpDesk {}
```

`health.enabled`, the two paths and `readyz.enabled` work the same on a platform that runs the fetch handler. Where `/readyz` is off by default, an edge runtime or a serverless deployment, turn it on with `readyz: { enabled: true }`, or rely on the platform's own checks.

---

## Troubleshooting

### `observability config is set but @frontmcp/observability is not installed`

The server has `observability` in its options, but the package isn't installed, so nothing is traced. The server works otherwise. Install it with `npm install @frontmcp/observability`. The Playground can't load the package, so an example with `observability` shows the same kind of warning in its **Logs** tab:

```ts main.ts active
import { App, FrontMcp, LogRecord, LogTransport, LogTransportInterface, Tool, ToolContext, z } from "@frontmcp/sdk";

export const records: LogRecord[] = [];

@LogTransport({ name: "MemoryTransport", description: "Keeps log records in memory, for tests" })
class MemoryTransport extends LogTransportInterface {
  log(record: LogRecord) {
    records.push(record);
  }
}

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string() } })
class CloseTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, status: "closed" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [CloseTicket] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: true,
  logging: { transports: [MemoryTransport] },
})
export default class Server {}
```

```ts missing.test.ts
import { test, expect } from "@frontmcp/testing";
import { records } from "./main";

test("the server warns, and calls still work", async ({ mcp }) => {
  const result = await mcp.tools.call("close_ticket", { id: "T-1" });
  expect(result.json()).toEqual({ id: "T-1", status: "closed" });
  const warning = records.find((r) => r.message.startsWith("observability config is set"));
  expect(warning?.levelName).toBe("warn");
  expect(warning?.message).toContain("@frontmcp/observability");
});
```

### `observability: tracing enabled but no TracerProvider configured`

No tracer provider was registered when the server started, so every span is a no-op. Register one before the server is built, as in [`observability`](#observability), or call `setupOTel()`. Setting `OTEL_EXPORTER_OTLP_ENDPOINT`, as the warning suggests, isn't enough on its own. (Before 1.9.2 the warning appeared even with a provider registered.)

### `ObservabilityPlugin is not installed or tracing is disabled`

The code read `this.telemetry` on a server with `observability: { tracing: false }`. The call fails with `TOOL_EXECUTION_ERROR`. Turn tracing on, or don't use `this.telemetry` there. Without `@frontmcp/observability` installed at all, `this.telemetry` is `undefined` instead, and the call fails with `Cannot read properties of undefined`.

### A failed span says `Internal FrontMCP error`

The span records what the client got in production, and for an error that isn't public that's FrontMCP's generic message, with an error id. The error's own message is in the `exception` event's stack, and in the server's error log under the same id: search the log for the id. Fail with a [`PublicMcpError`](https://frontmcp.dev/reference/sdk/fail) when the message is safe to show, and the span and the [request log](#request-logs) carry it. (Before 1.9.2, a failed tool's span wasn't exported at all, and `tools/call` ended with status OK. In 1.9.2, the status message was the error's own message, and empty after `this.fail()`.)

On FrontMCP 1.9.3, in a project that runs as ES modules, with `"type": "module"` in its `package.json`, a `PublicMcpError` was recorded this way too: the span said `Internal FrontMCP error` with `SERVER_ERROR`, the request log's `error` was a `GenericServerError`, and its error id wasn't the one the client got. After `this.fail()`, the stack was FrontMCP's, without the error's message. FrontMCP loads `@frontmcp/observability` with `require()`, which loads a second copy of `@frontmcp/sdk`, and that copy didn't recognize your code's error classes. Update to 1.9.4, where an ES module project records what the client got, as a CommonJS one does. This was checked in Node with FrontMCP 1.9.4, in an ES module project and a CommonJS one, with an in-memory span exporter and `requestLogs`.

### Structured logging writes nothing

`observability: { logging: true }` has no sinks. List at least one, like `{ sinks: [{ type: "stdout" }] }`.

### `Cannot find module '@frontmcp/observability'` at startup

`metrics.enabled` is `true` and the package isn't installed. `/metrics` needs it, and the server stops rather than start without the endpoint. Install `@frontmcp/observability`. `Cannot find module '@opentelemetry/sdk-trace-base'` means the package is there but that peer isn't: install it too.

### `metrics.auth is "token" but env var "…" is not set`

`MetricsTokenNotConfiguredError`: `auth: "token"` reads its token from `FRONTMCP_METRICS_TOKEN`, or the variable named in `tokenEnv`, when the server starts, and it isn't set. The server stops rather than serve metrics without the token. Set the variable, or use `auth: "public"` where the endpoint can't be reached from outside.

### `metrics.path "…" conflicts with an MCP transport route`

`MetricsPathConflictError`: `path` is `/mcp`, `/sse` or `/messages`, or under one of them, like `/mcp/metrics`. Pick another, like `/internal/metrics`.

### `/metrics` answers `404`

Check that `metrics.enabled` is `true` and that you're requesting `metrics.path`. Before 1.9.0, [`createFetchHandler()`](https://frontmcp.dev/reference/sdk/create-fetch-handler) didn't serve it at all.

### `/readyz` answers `503`

A probe returned `unhealthy`, threw, or took longer than `readyz.timeoutMs`. In development the `probes` field of the response says which, with its `error`. In production it's left out unless `includeDetails` is `true`.
