Observability and telemetry

Three things let you watch a FrontMCP server in production. Traces: with @frontmcp/observability installed, @FrontMcp({ observability }) records an OpenTelemetry span for each request, each call and each outbound this.fetch(), and this.telemetry adds your own. Metrics: @FrontMcp({ metrics }) serves process gauges and your counters at /metrics, for Prometheus, and a meter provider you register sends the same counters to an OpenTelemetry collector. Health checks: /healthz and /readyz tell an orchestrator whether to restart the server or send it traffic. Tracing and metrics send data to systems outside the process, so most of this page runs in your project, not in the Playground; each section says what the Playground can show.

@FrontMcp({
  info, apps,
  observability: true,          // or { tracing, logging, requestLogs }
  metrics: { enabled: true },   // GET /metrics
  health: { probes },           // GET /healthz, /readyz
})

Reference

observability

observability needs the @frontmcp/observability package, an optional peer of @frontmcp/sdk, and it records spans through whatever OpenTelemetry tracer provider you register. FrontMCP doesn't register one for you.

npm install @frontmcp/observability @opentelemetry/api @opentelemetry/sdk-trace-base @opentelemetry/exporter-trace-otlp-http
main.ts
import "reflect-metadata";
import { trace } from "@opentelemetry/api";
import { BasicTracerProvider, BatchSpanProcessor } from "@opentelemetry/sdk-trace-base";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { FrontMcp } from "@frontmcp/sdk";
import { HelpDeskApp } from "./help-desk.app";

// Any OpenTelemetry setup that registers a global tracer provider works. This one sends spans to a collector.
trace.setGlobalTracerProvider(
  new BasicTracerProvider({
    spanProcessors: [new BatchSpanProcessor(new OTLPTraceExporter({ url: "http://localhost:4318/v1/traces" }))],
  }),
);

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: true,
})
export default class Server {}

Register the provider before the server handles requests. Sending spans to a collector needs one running, so it can't be shown here; everything below about spans was checked by registering a provider that keeps them in memory.

@frontmcp/observability also exports setupOTel({ serviceName?, exporter?, endpoint?, serviceVersion? }), which registers a provider through OpenTelemetry's Node SDK and returns an async function that shuts it down. exporter is "console" (the default) or "otlp", which sends to endpoint plus /v1/traces; endpoint defaults to OTEL_EXPORTER_OTLP_ENDPOINT, then http://localhost:4318, and serviceName to OTEL_SERVICE_NAME, then "frontmcp-server". It needs @opentelemetry/sdk-node, @opentelemetry/resources and @opentelemetry/semantic-conventions, and @opentelemetry/exporter-trace-otlp-http for "otlp"; without them it throws setupOTel: @opentelemetry/sdk-node is required. Only those errors were checked here. The environment variables do nothing by themselves: they're read by setupOTel() and the otlp log sink, so setting them without calling setupOTel() exports nothing.

OptionTypeDefaultDescription
tracingboolean | TracingOptionstrueSpans for requests and calls, and this.telemetry. observability: true is { tracing: true }.
loggingboolean | { sinks, redactFields, includeStacks, staticFields }falseStructured log entries: each line of the server's log as an object, sent to the sinks you list.
requestLogsboolean | { maxEntries, includeSummaries, onRequestComplete }falseGathers each HTTP request's log lines and outcome into one object, and passes it to onRequestComplete. See request logs.

Without the package, the server starts anyway and writes one warning: observability config is set but @frontmcp/observability is not installed.

TracingOptions

OptionDefaultWhat it does
httpSpanstrueA server span for each HTTP request, named after it: POST /.
executionSpanstrueA span for each call (tools/call, resources/read, prompts/get), and one inside it for the tool, resource or prompt that ran.
fetchSpanstrueA client span for each this.fetch(), named after the method: GET. The traceparent that this.fetch() sends names this span; with false, it names the tool's span.
transportSpanstrueOn FrontMCP's HTTP server, a transport streamable-http span for each message of a session opened by a client on a protocol version before 2026-07-28.
authSpanstrueMeant for authentication flows. A public or static-token server records none.
flowStageEventstrueThe stage.* events on the request and call spans, like stage.findTool. false leaves them out.
hookSpansfalseA hook <stage> span, like hook willExecute, for each hook a plugin, app or entry runs, inside its call's span. Its attributes are frontmcp.hook.stage, frontmcp.hook.owner (the plugin's class name) and frontmcp.flow.name.
oauthSpanstrueA span for each request to an OAuth endpoint, named after it: oauth/token.
elicitationSpanstrueSpans for elicitation requests and results. Not checked here.
startupReporttrueA frontmcp.startup span once the server is ready, with frontmcp.startup.tools_count, resources_count, prompts_count, plugins_count and duration_ms.

Changed in 1.9.3: before, flowStageEvents, hookSpans, oauthSpans, elicitationSpans and startupReport were accepted but not read.

Spans

A 2026-07-28 tools/call records these spans. Each span's trace id is the request's this.context.traceContext.traceId, so a client that sends traceparent sees them in its own trace.

SpanKindParentAttributesEvents
POST /serverThe caller's span, from traceparenthttp.request.method, url.path, http.response.status_code, frontmcp.request.id, frontmcp.scope.id, mcp.session.idstage.traceRequest, stage.checkAuthorization, stage.router, stage.finalize, …
tools/callserverPOST /rpc.system ("mcp"), rpc.method, mcp.method.name, mcp.session.id, auth.authorities.resultstage.parseInput, stage.findTool, stage.validateInput, stage.validateOutput, stage.finalize, …
tool close_ticketinternaltools/callfrontmcp.tool.name, mcp.component.type ("tool"), mcp.component.key ("tool:close_ticket"), enduser.id, enduser.scopestage.execute.start, stage.execute.done, and yours
GETclienttools/callhttp.request.method, url.full, http.response.status_code

A resource read gives resources/read and resource tickets://T-1; a prompt, prompts/get and prompt triage; a tools/list, one tools/list span. Each stage event has a duration_ms attribute. mcp.session.id is a hash of the session id, never the id itself. The traceparent that this.fetch() sends names its GET span, so the service you call records its work under it. (Changed in 1.9.3: before, tools/call was a sibling of POST /, and this.fetch() passed on the caller's traceparent unchanged.)

Caveats

  • A failed call ends its spans with status ERROR, with what the client would get in production, whether the error was thrown or passed to this.fail(). The tool … span and the tools/call span both get the public message as their status message, and an exception event whose type is the error's code: no such ticket and PUBLIC_ERROR for a PublicMcpError, and for any other error, FrontMCP's Internal FrontMCP error. Please contact support with error ID: err_…, with the error id the client got, and TOOL_EXECUTION_ERROR when the tool threw it or SERVER_ERROR when it was passed to this.fail(). The event's stack still has the error's own message. See A failed span says Internal FrontMCP error. (Changed in 1.9.3: before, the status message was the error's own message, and empty after this.fail().)
  • ES module projects record the same. Changed in 1.9.4: in a project with "type": "module", every failure was recorded as SERVER_ERROR, a PublicMcpError too, with an error id the client didn't get: see the troubleshooting entry.
  • When no tracer provider is registered as the server starts, and NODE_ENV isn't development, FrontMCP logs observability: tracing enabled but no TracerProvider configured. In development it registers a provider that prints each span to the console instead, one line per span: ↳ SPAN tools/call [04bac610] 2.5ms SERVER rpc.system=mcp rpc.method=tools/call …. (Changed in 1.9.3: before, that failed silently on @opentelemetry/sdk-trace-base 2.x.)

this.telemetry

With observability on, tools, resources and prompts have this.telemetry, for spans and events of your own. What you add lands in the same trace, under the span of the tool, resource or prompt that's running.

search.tool.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";
import { TicketIndex } from "./ticket-index";

@Tool({ name: "search_tickets", description: "Search tickets by words in their title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    this.telemetry.addEvent("query-parsed", { terms: query.split(" ").length });
    const tickets = await this.telemetry.withSpan("ticket-index.search", async (span) => {
      span.setAttribute("db.query", query);
      return this.get(TicketIndex).search(query);
    });
    this.telemetry.setAttributes({ "results.count": tickets.length });
    return { tickets };
  }
}

The call records a ticket-index.search span, with db.query, as a child of tool search_tickets, and the tool's span gets the query-parsed event and the results.count attribute.

MemberDescription
withSpan(name, fn, attributes?)Runs fn(span) inside a new child span and ends it. If fn throws, the span ends with status ERROR and an exception event, and the error is rethrown.
startSpan(name, attributes?)Starts a child span and returns a TelemetrySpan. Call end() or endWithError() yourself: a span that isn't ended is never exported.
addEvent(name, attributes?)Adds an event to the running tool's, resource's or prompt's span.
setAttributes(attributes)Sets attributes on that span.
createCounter(name, description?)A counter for /metrics and for an OpenTelemetry collector.
traceIdThe request's trace id, the same as this.context.traceContext.traceId.
sessionIdA 16-character hash of the session id, the value of mcp.session.id.

Spans you start have frontmcp.request.id, frontmcp.scope.id and mcp.session.id besides your own attributes. Attribute values are strings, numbers or booleans.

TelemetrySpan memberDescription
setAttribute(key, value), setAttributes(attributes)Set attributes. Return the span, for chaining.
addEvent(name, attributes?)Add an event.
recordError(error)Record an exception event and set status ERROR. A later end() keeps the ERROR.
end()End the span, with status OK unless an error was recorded.
endWithError(error)End the span with status ERROR.
rawThe OpenTelemetry Span, for anything else.

this.telemetry exists only while tracing is on. Its type comes from @frontmcp/observability: TypeScript knows it once a file in your project imports the package, for createCounter() or with a bare import "@frontmcp/observability";. That includes prompts: PromptContext has it too.

It needs @frontmcp/observability, which the Playground can't load, so the examples of it on this page are code to run in your project.

Caveats

Structured logs

observability.logging turns each line of the server's log into an object with the request's ids, and sends it to sinks. It adds to the console and your transports; it doesn't replace them. Like the rest of observability, it doesn't run in the Playground; for a format of your own that does, write a log transport.

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: {
    logging: {
      sinks: [{ type: "stdout" }],
      redactFields: ["password", "token", "authorization"],
      staticFields: { service: "help-desk", region: "eu-west-1" },
    },
  },
  logging: { enableConsole: false },
})
export default class Server {}
{"service":"help-desk","region":"eu-west-1","timestamp":"2026-10-06T13:52:55.409Z","level":"info","severity_number":9,"message":"closing T-1","prefix":"tool:close_ticket","trace_id":"0af7651916cd43dd8448eb211c80319c","span_id":"b7ad6b7169203331","trace_flags":1,"request_id":"8f691bc7-290c-4f3f-ace9-0deadc6fc4fe","session_id_hash":"bef6cb302a57","scope_id":"root","flow_name":"tools:call-tool","elapsed_ms":10}
OptionDefaultDescription
sinksnoneWhere entries go. Required: logging: true alone has no sinks and writes nothing.
redactFields[]Keys to replace with "[REDACTED]" in the objects passed to the logger, in nested objects too, ignoring case.
includeStackstrueInclude the stack of an Error passed to the logger.
staticFields{}Fields added to every entry.
SinkWrites
{ type: "stdout" }One JSON object per line to stdout, or to stderr for a stdio server.
{ type: "console" }A readable line: [10:41:05.879] ERROR [716c9764:a556272f] [tool:pay] charge failed { orderId=O-1 } [PublicMcpError: card declined] (6ms), with the trace and request ids.
{ type: "callback", fn }Calls fn(entry).
{ type: "otlp", endpoint?, headers?, batchSize?, flushIntervalMs?, serviceName? }OTLP over HTTP to a collector; endpoint defaults to OTEL_EXPORTER_OTLP_ENDPOINT, then http://localhost:4318. Not checked here.
{ type: "winston", logger }, { type: "pino", logger }Your winston or pino logger. Not checked here.

Each entry has timestamp, level (lowercase), severity_number (OpenTelemetry's: 5, 6, 9, 13 or 17), message and prefix. Lines written during a request add trace_id, span_id (the caller's span), trace_flags, request_id, session_id_hash and scope_id, flow_name, the flow that wrote the line (http:request, session:verify, tools:call-tool, …), and elapsed_ms since the request started. Objects passed after the message are merged into attributes; an Error becomes error: { type, message, error_id, stack }.

Request logs

observability: { requestLogs } gathers what happened during each HTTP request into one object, and hands it to onRequestComplete once the response is ready: one per request, failed ones included. It needs @frontmcp/observability, and works whether or not observability.logging has sinks.

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: {
    requestLogs: {
      onRequestComplete: (log) => {
        if (log.status === "error") alerts.send(log); // your own code
      },
    },
  },
})
export default class Server {}
OptionDefaultDescription
onRequestCompleteCalled with the request's log. A promise it returns is awaited, and an error it throws is ignored.
maxEntries500Log lines kept per request. Later ones are dropped.
includeSummariesAccepted, but not read by FrontMCP 1.9.4.
FieldWhat it holds
request_id, trace_id, session_id_hash, scope_idThe request's ids, as on structured log entries.
start_time, end_time, duration_msWhen it ran.
http_method, http_path, rpc_method"POST", "/", "tools/call".
tool_name, resource_uri, prompt_nameWhat it called, read or got.
status, status_code"ok" or "error", and the HTTP status the server answered with: 200 for a tool call that failed too, 401 for a caller it refused. (Changed in 1.9.3: before, status_code was only set from 400.)
errorFor a failed call, { type, message, code, error_id }, as the client gets it in production, whether the error was thrown or passed to this.fail(): { type: "PublicMcpError", message: "no such ticket", code: "PUBLIC_ERROR", error_id: "err_…" } for a public error, and for any other, { type: "ToolExecutionError", message: "Internal FrontMCP error. Please contact support with error ID: err_…", code: "TOOL_EXECUTION_ERROR", error_id: "err_…" } when a tool threw it, or GenericServerError and SERVER_ERROR when it was passed to this.fail(). error_id is the errorId the client got. (Changed in 1.9.3: before, after this.fail() it was { type: "Error", message: "" }, and another error's own message was recorded.)
entriesEach line written during the request: { timestamp, level, message, stage, elapsed_ms, attributes? }, where stage is the flow, as flow_name above.
authenticated, auth_typetrue and "bearer" for a caller the server verified with a token; false and "anonymous" for an anonymous caller, or one it refused. (Changed in 1.9.3: before, authenticated was always false and auth_type missing.)
hooks_triggeredEach hook a plugin, app or entry ran during the request, as <flow>:<stage>: ["tools:call-tool:willExecute"]. (Changed in 1.9.3: before, it was always [].)

This was checked in Node, with createFetchHandler(); the Playground can't load @frontmcp/observability.

metrics

metrics adds a /metrics endpoint to FrontMCP's own HTTP server, and, since 1.9.0, to createFetchHandler(). It's off by default, and needs @frontmcp/observability with @opentelemetry/sdk-trace-base: without them, the server doesn't start (Cannot find module '@frontmcp/observability', or '@opentelemetry/sdk-trace-base').

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  metrics: { enabled: true, auth: "token" }, // reads FRONTMCP_METRICS_TOKEN
})
export default class Server {}
OptionTypeDefaultDescription
enabledbooleanfalseServe the endpoint.
pathstring"/metrics"Where. It can't be /mcp, /sse or /messages, or start with one of them: the server doesn't start, with MetricsPathConflictError.
format"prometheus" | "json""prometheus"Prometheus text, or { counters, gauges } as JSON.
auth"public" | "token" | { token }"public""token" requires Authorization: Bearer <token>, with the token read from tokenEnv at startup. { token: "…" } puts the token in the code, which is fine for local testing only.
tokenEnvstring"FRONTMCP_METRICS_TOKEN"The environment variable that holds the token. If it isn't set, the server doesn't start: MetricsTokenNotConfiguredError.
includeMetricsCategory[]every categoryOnly metrics whose names start with the categories' prefixes: process (frontmcp_process_, frontmcp_nodejs_), tools (frontmcp_tool_), resources (frontmcp_resource_), http (frontmcp_http_), storage (frontmcp_storage_), skills (frontmcp_skills_), auth (frontmcp_auth_), sessions (frontmcp_session_).
process{ eventLoopLag?, fdCount?, activeHandles? }all trueWhich Node gauges to collect. fdCount is Linux only.
RequestResponse
GET /metrics200, Content-Type: text/plain; charset=utf-8; version=0.0.4 (or application/json), Cache-Control: no-store.
Without Authorization, with auth: "token"401 { "error": "unauthorized", "message": "Missing or malformed Authorization header" }
With the wrong token403 { "error": "forbidden", "message": "Bearer token did not match the configured metrics token" }

The scrape holds the process gauges and your counters:

# HELP frontmcp_process_resident_memory_bytes Resident memory size in bytes
# TYPE frontmcp_process_resident_memory_bytes gauge
frontmcp_process_resident_memory_bytes 158318592
# HELP frontmcp_process_cpu_seconds_total CPU time consumed since collector start, by mode (seconds)
# TYPE frontmcp_process_cpu_seconds_total gauge
frontmcp_process_cpu_seconds_total{mode="user"} 0.022128
# TYPE frontmcp_nodejs_eventloop_lag_seconds gauge
frontmcp_nodejs_eventloop_lag_seconds{quantile="p99"} 5.11e-7
# TYPE helpdesk_tickets_closed_total counter
helpdesk_tickets_closed_total{priority="high"} 2

The gauges are frontmcp_process_resident_memory_bytes, _heap_bytes, _heap_used_bytes, _external_bytes, _uptime_seconds and _cpu_seconds_total{mode}, and frontmcp_nodejs_eventloop_lag_seconds{quantile="p99"}, _active_handles, _active_requests and, on Linux, _open_fds.

Skilled OpenAPI's counters

Skilled OpenAPI counts its bundles: frontmcp_skills_bundle_pulls_total (status, source) each time it applies one, and, when it checks a bundle's signature, frontmcp_skills_signature_verifications_total (status) and frontmcp_skills_signature_failures_total (reason, like missing_integrity). Each bundle it applies also records a skill.bundle.swap span, with bundle_id, version and skill_count. It does so only when ObservabilityPlugin from @frontmcp/observability comes before it in plugins:

main.ts
import { ObservabilityPlugin } from "@frontmcp/observability";
import { SkilledOpenApiPlugin } from "@frontmcp/plugin-skilled-openapi";

@FrontMcp({
  info: { name: "billing", version: "1.0.0" },
  apps: [],
  metrics: { enabled: true },
  plugins: [ObservabilityPlugin.init({ tracing: true }), SkilledOpenApiPlugin.init({ source, trustedKeys })],
})
export default class Server {}
# TYPE frontmcp_skills_bundle_pulls_total counter
frontmcp_skills_bundle_pulls_total{source="unknown",status="ok"} 1

With observability: true instead, or with ObservabilityPlugin after it, the plugin can't reach the counters and records nothing, though tracing works. This was checked in Node with createFetchHandler(), an inline bundle and an in-memory span exporter.

Caveats

  • FrontMCP 1.9.4 counts no requests, tool calls or errors. The scrape is the process gauges and the counters you create, plus Skilled OpenAPI's own when it's installed, and only then: see below. The include categories other than process and skills only match counters you name with their prefixes.
  • Every server FrontMCP builds serves it: the one @FrontMcp and FrontMcpInstance.bootstrap() start, the handler FrontMcpInstance.createHandler() returns for serverless platforms, and createFetchHandler()'s, which before 1.9.0 ignored metrics and answered /metrics with 404. The Playground can't load @frontmcp/observability, so it can't show it.
  • A counter's description isn't printed as a # HELP line.

health

FrontMCP answers health checks without any configuration. health adds your own checks, or changes the paths. The server's other options are on @FrontMcp.

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  health: {
    probes: [
      {
        name: "tickets-db",
        async check() {
          const started = Date.now();
          await db.query("select 1");
          return { status: "healthy", latencyMs: Date.now() - started };
        },
      },
    ],
    readyz: { timeoutMs: 2000 },
  },
})
export default class Server {}
OptionTypeDefaultDescription
enabledbooleantruefalse removes /healthz, /readyz and /health.
healthzPathstring"/healthz"The liveness path. /health answers the same whatever this is.
readyzPathstring"/readyz"The readiness path.
probes{ name, check() }[][]Your checks, run in parallel on every /readyz. check() returns { status, latencyMs?, details?, error? }, with status "healthy", "degraded" or "unhealthy".
includeDetailsbooleantrue in development, false in productionInclude each probe's result in /readyz.
readyz{ enabled?, timeoutMs? }on, unless the runtime is an edge one or the deployment is serverless; timeoutMs: 5000enabled: false removes /readyz: it's a 404. A probe that takes longer than timeoutMs is unhealthy, with "Probe timed out after 2000ms".
RequestFrontMCP's HTTP server (and createHandler()) answers
GET /healthz, GET /health200 { "status": "ok", "server": { "name", "version" }, "runtime": { "platform", "runtime", "deployment", "env" }, "uptime": 3600.5 }. No checks: if the process can answer, it's alive.
GET /readyz200 { "status": "ready", "totalLatencyMs", "catalog": { "toolsHash", "toolCount", "resourceCount", "promptCount", "skillCount", "agentCount" }, "probes": { … } }, or 503 with "status": "not_ready" when a probe is unhealthy, throws (its message becomes error) or times out. degraded still counts as ready.

The probes it lists include FrontMCP's own session-store check next to yours, and a remote:<app id> check for each remote app, which is degraded, and so still ready, until the first check of the remote has run. toolsHash is a hash of the tool names, the same on every copy of a server with the same tools, so a replica that's out of step stands out.

createFetchHandler() honours health too: enabled, both paths, probes, includeDetails and readyz. Its /readyz answers as above, with "transport": "web-fetch" added, and 503 when a probe is unhealthy. Its /healthz, and /health, answer the shorter 200 { "status": "ok", "server": { "name", "version" }, "transport": "web-fetch" }, with no runtime or uptime. That's what the Playground runs, so the example below shows it. (Before 1.8.7 it ignored health: /healthz and /readyz both returned a fixed 200, no probe ran, and /health was a 404.)

Trace context helpers

@frontmcp/sdk exports the functions it uses to read W3C trace context. They work without @frontmcp/observability. Each returns { traceId, parentId, traceFlags, raw }, where raw is a traceparent value.

FunctionReturns
parseTraceContext(headers)The trace context of a request's headers: from traceparent, else from x-frontmcp-trace-id (32 hex characters) with a new parent id, else a new trace. A malformed or all-zero traceparent counts as missing.
generateTraceContext()A new trace, sampled.
createChildSpanContext(parent)The same trace with a new parent id: the traceparent to send for work you do on the trace's behalf.

Usage

Following a request across services

A request's trace id comes from the caller. A traceparent header continues the caller's trace; without one, FrontMCP also accepts an x-frontmcp-trace-id header, and without either, the request starts a trace of its own. The same id is in this.context.traceContext, on every line of this.contextLogger (its first eight characters), in the traceparent that this.fetch() sends on, and, with tracing on, in every span. That's what lets you go from a log line to the trace, and from the trace to the services it called. This part runs in the Playground:

Open
import { App, Tool, ToolContext } from "@frontmcp/sdk";

@Tool({ name: "trace_info", description: "Show this request's trace. For debugging.", inputSchema: {} })
export class TraceInfo extends ToolContext {
  async execute() {
    this.contextLogger.info("looking up the trace");
    const { traceId, parentId, traceFlags } = this.context.traceContext;
    return { traceId, parentId, traceFlags };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [TraceInfo] })
export class HelpDesk {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The Logs tab shows the tagged line: its second id is the start of the traceId in the result. A traceparent inside the request's _meta isn't read; only the header is (see Joining a trace). Work outside a request, like a job, has no trace of its own, and this.fetch() sends no traceparent there: start one with generateTraceContext() and send its raw yourself.

Adding your own spans

The spans FrontMCP records show which call was slow. Your own spans show which part of it: a database query, a call to a model, a loop over pages of results. Wrap each in this.telemetry.withSpan(), and record what a later reader will want to filter by as attributes:

close-ticket.tool.ts
import { PublicMcpError, Tool, ToolContext, z } from "@frontmcp/sdk";
import { Tickets } from "./tickets";

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string() } })
export class CloseTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    const ticket = await this.telemetry.withSpan("tickets.load", () => this.get(Tickets).load(id), { "ticket.id": id });
    if (!ticket) this.fail(new PublicMcpError(`There's no ticket ${id}.`));

    this.telemetry.setAttributes({ "ticket.priority": ticket.priority });
    await this.telemetry.withSpan("tickets.close", async (span) => {
      await this.get(Tickets).close(id);
      span.addEvent("customer-notified");
    });
    return { id, status: "closed" };
  }
}

A call gives tool close_ticket, with the ticket.priority attribute, and inside it tickets.load and tickets.close, each timed. If load() throws, tickets.load ends with status ERROR and the error as an exception event, and the call fails as usual. This needs @frontmcp/observability, so it doesn't run in the Playground.

Checking spans in a test

@frontmcp/observability has helpers for tests: createTestTracer() gives a tracer provider that keeps spans in memory, and assertSpanExists(), assertSpanAttribute(), findSpan() and findSpansByAttribute() search them. Register the provider globally, or FrontMCP's spans don't reach it. The server must run in the test's process, as it does with createFetchHandler():

close-ticket.spans.test.ts
import { trace } from "@opentelemetry/api";
import { assertSpanAttribute, assertSpanExists, createTestTracer } from "@frontmcp/observability";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDeskApp } from "./help-desk.app";

const { provider, exporter, cleanup } = createTestTracer();
trace.setGlobalTracerProvider(provider); // without this, the exporter stays empty

afterEach(() => exporter.reset());
afterAll(() => cleanup());

test("closing a ticket records a tickets.close span", async () => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDeskApp],
    observability: true,
  });
  await handler(closeTicketRequest("T-1")); // a 2026-07-28 tools/call request

  const spans = exporter.getFinishedSpans();
  assertSpanExists(spans, "tickets.close");
  assertSpanAttribute(assertSpanExists(spans, "tool close_ticket"), "mcp.component.key", "tool:close_ticket");
});

assertSpanExists() returns the span, or throws Expected span "tickets.close" but found: [tool close_ticket, tools/call, POST /]. This is a test for your own project's runner; the Playground's tests can't import @frontmcp/observability.

Counting things

A counter adds up across requests, which spans can't: tickets closed, cache hits, calls to a paid API. Create it once, with a name in snake case ending in _total, and add to it with inc(by?, attributes?). /metrics shows every counter, one line per set of attributes:

close-ticket.tool.ts
import { createCounter } from "@frontmcp/observability";
import { Tool, ToolContext, z } from "@frontmcp/sdk";

const ticketsClosed = createCounter("helpdesk_tickets_closed_total", "Tickets closed");

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string(), priority: z.enum(["low", "high"]) } })
export class CloseTicket extends ToolContext {
  async execute({ id, priority }: { id: string; priority: "low" | "high" }) {
    ticketsClosed.inc(1, { priority });
    return { id, status: "closed" };
  }
}
# TYPE helpdesk_tickets_closed_total counter
helpdesk_tickets_closed_total{priority="high"} 2

this.telemetry.createCounter(name) returns the same counter. Keep attribute values to a small, fixed set, like a priority or a status: each new value is a new series to store. Counts live in the process's memory, so each replica reports its own, and they start at zero on restart. getMetricSnapshot() and getCounterTotal(name) from @frontmcp/observability read them, which is handy in tests. This needs the package, so it doesn't run in the Playground.

Sending counters to an OpenTelemetry collector

/metrics waits for Prometheus to scrape it. To push the same counters to an OpenTelemetry collector, register a meter provider: every counter also reports to OpenTelemetry's global meter, as the instrumentation scope @frontmcp/observability. Neither observability nor metrics registers one, since observability sets up tracing only, and neither option is needed: the counters come from @frontmcp/observability alone. Register the provider in a file of its own, and import it first:

npm install @opentelemetry/api @opentelemetry/sdk-metrics @opentelemetry/exporter-metrics-otlp-http @opentelemetry/resources
telemetry.ts
import { metrics } from "@opentelemetry/api";
import { OTLPMetricExporter } from "@opentelemetry/exporter-metrics-otlp-http";
import { resourceFromAttributes } from "@opentelemetry/resources";
import { MeterProvider, PeriodicExportingMetricReader } from "@opentelemetry/sdk-metrics";

export const meterProvider = new MeterProvider({
  resource: resourceFromAttributes({ "service.name": "help-desk" }),
  readers: [
    new PeriodicExportingMetricReader({
      exporter: new OTLPMetricExporter({ url: "http://localhost:4318/v1/metrics" }),
      exportIntervalMillis: 10_000,
    }),
  ],
});

metrics.setGlobalMeterProvider(meterProvider);
main.ts
import "reflect-metadata";
import "./telemetry"; // before any file that creates a counter
import { FrontMcp } from "@frontmcp/sdk";
import { HelpDeskApp } from "./help-desk.app";

@FrontMcp({ info: { name: "help-desk", version: "1.0.0" }, apps: [HelpDeskApp] })
export default class Server {}

With the close_ticket tool from Counting things in HelpDeskApp, two calls with priority: "high" reached the collector as helpdesk_tickets_closed_total, a sum of 2 with the attribute priority: "high", from a resource whose service.name is help-desk. meterProvider.forceFlush() sends what has been counted so far, which a test can use instead of waiting for the interval.

The import order matters. createCounter() asks OpenTelemetry for the counter the first time it sees a name, and keeps it. A counter created before the provider is registered reports to a meter that does nothing, for as long as the process runs, while /metrics and getCounterTotal() still count it. A counter at the top of a tool's file is created when the file is imported, which is why telemetry.ts comes first. this.telemetry.createCounter(name) runs later, inside a call, but returns the counter the name already has.

setupOTel(), from the observability section, registers a meter provider too, through OpenTelemetry's Node SDK. With that SDK's defaults (checked with @opentelemetry/sdk-node 0.219), it sends the counters every 60 seconds, as OTLP over protobuf, to OTEL_EXPORTER_OTLP_ENDPOINT plus /v1/metrics, which is http://localhost:4318/v1/metrics when the variable isn't set. Its endpoint option moves the traces only. OTEL_METRICS_EXPORTER=none turns the meter provider off. Nothing reports that the collector isn't there.

This needs the package, so it doesn't run in the Playground. It was checked in Node with FrontMCP 1.9.4, with an in-memory exporter and with a small HTTP server standing in for the collector.

Answering health checks

An orchestrator asks two questions. /healthz: is the process alive, or should it be restarted? /readyz: can it take traffic now? Point liveness at the first and readiness at the second, and put what the server can't work without, like its database, in health.probes. A failing probe turns /readyz into a 503, and the orchestrator stops sending traffic until it passes, while /healthz stays 200.

A server behind createFetchHandler(), like this Playground's, does the same:

Open
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";

const unhealthy = { name: "tickets-db", check: async () => ({ status: "unhealthy" as const, error: "connection refused" }) };

const requestTo = async (path: string, health: object = { probes: [unhealthy] }) => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDesk],
    health: { includeDetails: true, ...health },
  });
  return handler(new Request(`https://desk.example.com${path}`));
};

test("/healthz is 200 whatever the probes say", async () => {
  const response = await requestTo("/healthz");
  expect(response.status).toBe(200);
  expect(await response.json()).toEqual({ status: "ok", server: { name: "help-desk", version: "1.0.0" }, transport: "web-fetch" });
});

test("/readyz runs `health.probes` and answers 503 while one is unhealthy", async () => {
  const response = await requestTo("/readyz");
  expect(response.status).toBe(503);
  const body = await response.json();
  expect(body.status).toBe("not_ready");
  expect(body.probes["tickets-db"]).toEqual({ status: "unhealthy", error: "connection refused" });
  expect(body.probes["session-store"].status).toBe("healthy");
});

test("/readyz is 200 once the probes pass", async () => {
  const healthy = { name: "tickets-db", check: async () => ({ status: "healthy" as const, latencyMs: 3 }) };
  const response = await requestTo("/readyz", { probes: [healthy] });
  expect(response.status).toBe(200);
  expect((await response.json()).status).toBe("ready");
});

test("a probe that takes longer than `readyz.timeoutMs` is unhealthy", async () => {
  const slow = { name: "slow", check: () => new Promise<{ status: "healthy" }>((resolve) => setTimeout(() => resolve({ status: "healthy" }), 300)) };
  const response = await requestTo("/readyz", { probes: [slow], readyz: { timeoutMs: 50 } });
  expect(response.status).toBe(503);
  expect((await response.json()).probes.slow.error).toBe("Probe timed out after 50ms");
});

test("`healthzPath` and `readyzPath` move the paths, and `enabled: false` removes them", async () => {
  const moved = { healthzPath: "/live", readyzPath: "/ready", probes: [] };
  expect((await requestTo("/live", moved)).status).toBe(200);
  expect((await requestTo("/ready", moved)).status).toBe(200);
  expect((await requestTo("/healthz", moved)).status).toBe(404);
  expect((await requestTo("/healthz", { enabled: false })).status).toBe(404);
});

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

health.enabled, the two paths and readyz.enabled work the same on a platform that runs the fetch handler. Where /readyz is off by default, an edge runtime or a serverless deployment, turn it on with readyz: { enabled: true }, or rely on the platform's own checks.


Troubleshooting

observability config is set but @frontmcp/observability is not installed

The server has observability in its options, but the package isn't installed, so nothing is traced. The server works otherwise. Install it with npm install @frontmcp/observability. The Playground can't load the package, so an example with observability shows the same kind of warning in its Logs tab:

Open
import { App, FrontMcp, LogRecord, LogTransport, LogTransportInterface, Tool, ToolContext, z } from "@frontmcp/sdk";

export const records: LogRecord[] = [];

@LogTransport({ name: "MemoryTransport", description: "Keeps log records in memory, for tests" })
class MemoryTransport extends LogTransportInterface {
  log(record: LogRecord) {
    records.push(record);
  }
}

@Tool({ name: "close_ticket", description: "Close a support ticket", inputSchema: { id: z.string() } })
class CloseTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, status: "closed" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [CloseTicket] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  observability: true,
  logging: { transports: [MemoryTransport] },
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

observability: tracing enabled but no TracerProvider configured

No tracer provider was registered when the server started, so every span is a no-op. Register one before the server is built, as in observability, or call setupOTel(). Setting OTEL_EXPORTER_OTLP_ENDPOINT, as the warning suggests, isn't enough on its own. (Before 1.9.2 the warning appeared even with a provider registered.)

ObservabilityPlugin is not installed or tracing is disabled

The code read this.telemetry on a server with observability: { tracing: false }. The call fails with TOOL_EXECUTION_ERROR. Turn tracing on, or don't use this.telemetry there. Without @frontmcp/observability installed at all, this.telemetry is undefined instead, and the call fails with Cannot read properties of undefined.

A failed span says Internal FrontMCP error

The span records what the client got in production, and for an error that isn't public that's FrontMCP's generic message, with an error id. The error's own message is in the exception event's stack, and in the server's error log under the same id: search the log for the id. Fail with a PublicMcpError when the message is safe to show, and the span and the request log carry it. (Before 1.9.2, a failed tool's span wasn't exported at all, and tools/call ended with status OK. In 1.9.2, the status message was the error's own message, and empty after this.fail().)

On FrontMCP 1.9.3, in a project that runs as ES modules, with "type": "module" in its package.json, a PublicMcpError was recorded this way too: the span said Internal FrontMCP error with SERVER_ERROR, the request log's error was a GenericServerError, and its error id wasn't the one the client got. After this.fail(), the stack was FrontMCP's, without the error's message. FrontMCP loads @frontmcp/observability with require(), which loads a second copy of @frontmcp/sdk, and that copy didn't recognize your code's error classes. Update to 1.9.4, where an ES module project records what the client got, as a CommonJS one does. This was checked in Node with FrontMCP 1.9.4, in an ES module project and a CommonJS one, with an in-memory span exporter and requestLogs.

Structured logging writes nothing

observability: { logging: true } has no sinks. List at least one, like { sinks: [{ type: "stdout" }] }.

Cannot find module '@frontmcp/observability' at startup

metrics.enabled is true and the package isn't installed. /metrics needs it, and the server stops rather than start without the endpoint. Install @frontmcp/observability. Cannot find module '@opentelemetry/sdk-trace-base' means the package is there but that peer isn't: install it too.

metrics.auth is "token" but env var "…" is not set

MetricsTokenNotConfiguredError: auth: "token" reads its token from FRONTMCP_METRICS_TOKEN, or the variable named in tokenEnv, when the server starts, and it isn't set. The server stops rather than serve metrics without the token. Set the variable, or use auth: "public" where the endpoint can't be reached from outside.

metrics.path "…" conflicts with an MCP transport route

MetricsPathConflictError: path is /mcp, /sse or /messages, or under one of them, like /mcp/metrics. Pick another, like /internal/metrics.

/metrics answers 404

Check that metrics.enabled is true and that you're requesting metrics.path. Before 1.9.0, createFetchHandler() didn't serve it at all.

/readyz answers 503

A probe returned unhealthy, threw, or took longer than readyz.timeoutMs. In development the probes field of the response says which, with its error. In production it's left out unless includeDetails is true.