# Health checks and metrics

> What /healthz, /readyz and /metrics answer on each FrontMCP runtime, and how to point container health checks, Kubernetes probes, load balancers and Prometheus at them.

Source: https://frontmcp.dev/reference/deployment/health-and-metrics

Orchestrators and load balancers ask a server two questions: is the process alive, and should it get traffic now? FrontMCP answers the first at `/healthz` and the second at `/readyz`, and serves process gauges and your counters at `/metrics` for Prometheus. What each path answers depends on how the server runs: FrontMCP's Node server runs readiness probes, a serverless function has no `/readyz`, and a Worker has `/healthz` and `/readyz` only when you turn the second on. This page is about wiring them into a deployment; [Observability and telemetry](https://frontmcp.dev/reference/server/observability#health) has every option of `health` and `metrics`.

```ts
@FrontMcp({ info, apps, health: { probes, readyz: { timeoutMs } }, metrics: { enabled: true, auth: "token" } })
```

---

## Reference

### What each runtime answers

| | Node server: `node`, `distributed`, `cli serve`, `frontmcp dev` | Serverless: `vercel`, `lambda` | Fetch handler: `cloudflare`, [`createFetchHandler()`](https://frontmcp.dev/reference/sdk/create-fetch-handler) |
| --- | --- | --- | --- |
| `GET /healthz` | `200` `{ status: "ok", server, runtime: { platform, runtime, deployment, env }, uptime }`. No checks. | The same, with `"deployment": "serverless"` | `200` `{ status: "ok", server, transport: "web-fetch" }`. No checks. |
| `GET /health` | The same as `/healthz` | The same as `/healthz` | The same as `/healthz` |
| `GET /readyz` | Runs every probe: `200` `"status": "ready"`, or `503` `"not_ready"` | `404`: off in serverless mode | Off on a Worker (`404`) until `health.readyz.enabled: true`; in `createFetchHandler()` on Node it's on. When on, it runs every probe like the Node server, with `"transport": "web-fetch"` in the answer. |
| `GET /metrics` | With `metrics.enabled` | Not checked for this page | With `metrics.enabled`, since 1.9, and on a Worker since 1.9.2, without the process gauges: see [Cloudflare Workers](https://frontmcp.dev/reference/deployment/cloudflare-workers#what-the-worker-serves). |
| `health` options | Applied | Not checked for this page | Applied: `enabled`, `healthzPath`, `readyzPath`, `probes`, `includeDetails`, `readyz` |
| `Cache-Control` | `no-cache, no-transform` | `no-cache, no-transform` | `no-store` |

`health.enabled: false` removes `/healthz`, `/health` and `/readyz` on every runtime. A custom `healthzPath` replaces `/healthz` only: `/health` still answers.

> **Note**
In 1.8.6 a Worker, or anything behind `createFetchHandler()`, ignored `health`: `enabled: false` and a custom path changed nothing, `/readyz` was the same fixed `200` as `/healthz`, no probe ran, and `/health` was a `404`. Now the options apply, `/readyz` runs probes and answers `503` when one fails, and `/health` is an alias of `/healthz`. On a Worker `/readyz` is still off unless you ask for it.

On every runtime:

- **Health checks don't need credentials.** With `static`, `transparent`, `local` or `remote` auth, `/healthz` and `/readyz` still answer `200` without a token, while an MCP request without one gets `401`.
- **They go through the host check.** A server that checks `Host` headers ([host checks](https://frontmcp.dev/reference/deployment/node#host-checks)) answers a health check sent to another host name, like a pod's IP address, with `403` `Invalid Host header`.
- **They don't see missing secrets.** A production server without `MCP_SESSION_SECRET` answers `/healthz` with `200`, and the `initialize` of clients before MCP 2026-07-28 with `500` and `{"error":"server_misconfigured","code":"SESSION_SECRET_REQUIRED",…}`. Only a Worker without `JWT_SECRET` fails its health checks too ([production secrets](https://frontmcp.dev/reference/sdk/create-fetch-handler#production-secrets)).
- **On a Worker, they fail when the server can't be built.** Every request, health checks included, answers `500 server_misconfigured` or `503 server_unavailable` with `Retry-After` until a retry builds it: see [Cloudflare Workers](https://frontmcp.dev/reference/deployment/cloudflare-workers#when-the-server-cant-be-built).

### `/readyz` on the Node server

It runs FrontMCP's own `session-store` probe, which pings Redis or SQLite when sessions are stored there, and the `probes` you add, all at once, each with `readyz.timeoutMs` (5 seconds by default). Any `unhealthy` probe, one that throws, or one that times out makes it `503`; `degraded` counts as ready.

```json
{"status":"ready","totalLatencyMs":0,"catalog":{"toolsHash":"…","toolCount":1,"resourceCount":0,"promptCount":0,"skillCount":0,"agentCount":0},"probes":{"session-store":{"status":"healthy","latencyMs":0},"tickets-db":{"status":"healthy","latencyMs":2}}}
```

Each [remote app](https://frontmcp.dev/reference/server/remote) adds a probe too, named `remote:<app id>` (the id is the app's name unless you gave one), like `remote:status`. It reports what FrontMCP's own check of the remote, every 30 seconds, last found. Before the first check it's `degraded` with `"state": "unknown"`, which counts as ready, so `/readyz` isn't `503` while a server with a remote app starts. Started against a remote that was up, the first `/readyz` after the server listened already showed it `healthy`.

In production the `probes` field is left out, so the status is all a caller learns; `includeDetails: true` puts it back. `catalog.toolsHash` is the same on every instance with the same tools: [High availability](https://frontmcp.dev/reference/deployment/high-availability#checking-that-replicas-serve-the-same-tools) compares it across replicas. When Redis is down, the `session-store` probe is `unhealthy` and `/readyz` answers `503` until it's back: also at startup, where FrontMCP now retries the connection with a growing delay instead of giving up for good. See [Redis](https://frontmcp.dev/reference/deployment/redis#when-redis-cant-be-reached).

### `/metrics`

`metrics: { enabled: true }` needs two packages, and the server doesn't start without either:

```bash
npm install @frontmcp/observability @opentelemetry/sdk-trace-base
```

Without `@opentelemetry/sdk-trace-base`, startup logs `Failed to wire /metrics endpoint Error: Cannot find module '@opentelemetry/sdk-trace-base'` and the process exits. With them, `GET /metrics` answers Prometheus text (`Content-Type: text/plain; charset=utf-8; version=0.0.4`, `Cache-Control: no-store`): process and Node gauges, and the counters you create. With `auth: "token"`, a request without `Authorization: Bearer <FRONTMCP_METRICS_TOKEN>` gets `401`, one with the wrong token `403`. [`metrics`](https://frontmcp.dev/reference/server/observability#metrics) has the options and the output.

`createFetchHandler()` serves `/metrics` the same way since 1.9, with the same `401` and `403`, and `Content-Type: text/plain; version=0.0.4; charset=utf-8`; in 1.8.7 it answered `404`. It needs the same two packages: without the second, `createFetchHandler()` rejects with `Cannot find module '@opentelemetry/sdk-trace-base'`. A Worker serves it too since 1.9.2, with `Cache-Control: no-store`, and only the counters: it has no process to measure, so it leaves the `frontmcp_process_*` gauges out (in 1.9.2 it listed them, all `0`).

#### Caveats

- `/healthz` checks nothing: an instance that can't reach its database still answers `200`. Put what the server can't work without in `health.probes`, and use `/readyz` for readiness.
- A `GET` to the MCP endpoint isn't a health check: without a session, FrontMCP's Node server answers it with Express's `404` page (`Cannot GET /`), and a fetch handler with `406 Not Acceptable` unless it sends `Accept: text/event-stream`.
- The Kubernetes and load-balancer settings below weren't run on Kubernetes. With FrontMCP 1.9.3: the Docker health check of the image `frontmcp create` writes, `/health` beside a custom `healthzPath`, and `/metrics` on a Worker and from `createFetchHandler()` in Node. With 1.9.2: the fetch handler's health paths with `createFetchHandler()` (the Playground below) and `wrangler dev`, and `/readyz` on the Node server. `/metrics` on the Node server was last run with 1.9.1, and the Prometheus scrape (Prometheus 3.5 in Docker, with the token in a credentials file) with 1.8.7.

---

## Usage

### Checking health from inside a container

`node:24-slim` has no `curl`, so the Dockerfile `frontmcp create` writes checks with Node itself:

```text title="ci/Dockerfile"
HEALTHCHECK --interval=30s --timeout=5s --start-period=15s --retries=3 \
  CMD node -e "fetch('http://127.0.0.1:' + (process.env.PORT || 3000) + '/health').then((r) => process.exit(r.ok ? 0 : 1), () => process.exit(1))"
```

`docker inspect` reported the container `healthy` after the first check. It asks `/health`, which stays put when `health.healthzPath` moves `/healthz`: with `healthzPath: "/live"`, `/health` and `/live` answered `200` and `/healthz` `404`. (In 1.9.2 the check asked `/healthz`.) The Nx `server` generator's Dockerfile has the same check since 1.9.3. With `FRONTMCP_ALLOWED_HOSTS` set, include `127.0.0.1:3000` in it, or this check gets `403` and the container is reported `unhealthy`. [Node.js and Docker](https://frontmcp.dev/reference/deployment/node#building-a-docker-image) has the whole Dockerfile.

### Pointing Kubernetes at the probes

Liveness at `/healthz`, readiness at `/readyz`, with the probe's timeout longer than `readyz.timeoutMs`:

```yaml deployment.yaml
containers:
  - name: help-desk
    image: registry.example.com/help-desk:1.2.0
    env:
      - name: FRONTMCP_BIND_ADDRESS
        value: all
      - name: FRONTMCP_ALLOWED_HOSTS
        value: desk.example.com
    ports:
      - containerPort: 3000
    livenessProbe:
      httpGet:
        path: /healthz
        port: 3000
        httpHeaders:
          - name: Host
            value: desk.example.com
      periodSeconds: 10
    readinessProbe:
      httpGet:
        path: /readyz
        port: 3000
        httpHeaders:
          - name: Host
            value: desk.example.com
      periodSeconds: 15
      timeoutSeconds: 6
```

The kubelet calls the pod by its IP address, which isn't in `FRONTMCP_ALLOWED_HOSTS`, so the probes send the public host name instead. Without `FRONTMCP_ALLOWED_HOSTS`, a server on all interfaces checks no host, and the headers aren't needed.

### Scraping metrics with Prometheus

```ts main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDesk],
  metrics: { enabled: true, auth: "token" }, // FRONTMCP_METRICS_TOKEN
})
export default class Server {}
```

```yaml prometheus.yml
scrape_configs:
  - job_name: help-desk
    metrics_path: /metrics
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/metrics-token
    static_configs:
      - targets: ["help-desk:3000"]
```

With the same token in `FRONTMCP_METRICS_TOKEN` and in the credentials file, Prometheus reports the target `up`, and `frontmcp_process_resident_memory_bytes{job="help-desk"}` returns the server's memory. `/metrics` is on the same port as MCP; keep it off the public internet at your proxy, or rely on the token.

### Health checks behind authentication, on a Worker

A load balancer or uptime monitor can check a Worker, or any server behind `createFetchHandler()`, without a token. The host check applies to them, and so do `health`'s options:

```ts health.test.ts active
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";

const config = {
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDesk],
  http: { entryPath: "/mcp" },
  auth: { mode: "static" as const, tokens: ["sk_desk_1"] },
  health: { readyz: { enabled: true } }, // off by default on a Worker
};

const get = async (handler: (request: Request) => Promise<Response>, path: string) => handler(new Request(`https://desk.example.com${path}`));

test("health checks need no token, though MCP calls do", async () => {
  const handler = await FrontMcpInstance.createFetchHandler(config);
  for (const path of ["/healthz", "/health"]) {
    const res = await get(handler, path);
    expect(res.status).toBe(200);
    expect(res.headers.get("cache-control")).toBe("no-store");
    expect(await res.json()).toEqual({ status: "ok", server: { name: "help-desk", version: "1.0.0" }, transport: "web-fetch" });
  }
  const ready = await get(handler, "/readyz");
  expect(ready.status).toBe(200);
  expect(await ready.json()).toMatchObject({ status: "ready", transport: "web-fetch", catalog: { toolCount: 1 } });
  const call = await handler(
    new Request("https://desk.example.com/mcp", {
      method: "POST",
      headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/list" },
      body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/list", params: { _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } } }),
    }),
  );
  expect(call.status).toBe(401);
});

test("a failing probe makes `/readyz` answer 503, and `/healthz` stays 200", async () => {
  let up = true;
  const probes = [{ name: "tickets-db", check: async () => (up ? { status: "healthy" as const } : { status: "unhealthy" as const, error: "no connection" }) }];
  const handler = await FrontMcpInstance.createFetchHandler({ ...config, health: { readyz: { enabled: true }, probes, includeDetails: true } });
  expect((await get(handler, "/readyz")).status).toBe(200);
  up = false;
  const down = await get(handler, "/readyz");
  expect(down.status).toBe(503);
  expect(await down.json()).toMatchObject({ status: "not_ready", probes: { "tickets-db": { status: "unhealthy", error: "no connection" } } });
  expect((await get(handler, "/healthz")).status).toBe(200);
});

test("`health` sets the paths, and `enabled: false` removes them", async () => {
  const renamed = await FrontMcpInstance.createFetchHandler({ ...config, health: { healthzPath: "/live", readyzPath: "/ready", readyz: { enabled: true } } });
  expect((await get(renamed, "/live")).status).toBe(200);
  expect((await get(renamed, "/ready")).status).toBe(200);
  expect((await get(renamed, "/healthz")).status).toBe(404);
  expect((await get(renamed, "/health")).status).toBe(200); // still an alias of the liveness path

  const off = await FrontMcpInstance.createFetchHandler({ ...config, health: { enabled: false } });
  for (const path of ["/healthz", "/health", "/readyz"]) expect((await get(off, path)).status).toBe(404);
});

test("a check sent to another host name is refused", async () => {
  const guarded = await FrontMcpInstance.createFetchHandler({
    ...config,
    http: { entryPath: "/mcp", security: { dnsRebindingProtection: { enabled: true, allowedHosts: ["desk.example.com"] } } },
  });
  expect((await guarded(new Request("http://10.0.3.7:8787/healthz"))).status).toBe(403);
  expect((await guarded(new Request("https://desk.example.com/healthz"))).status).toBe(200);
});
```

```ts help-desk.app.ts
import { App, Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "get_ticket", description: "Get one support ticket by its id", inputSchema: { id: z.string() } })
class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [GetTicket] })
export class HelpDesk {}
```

In production `/readyz` leaves `probes` out unless `includeDetails` is `true`, as on the Node server; the test sets it, so it passes with any `NODE_ENV`. A probe that doesn't answer within `health.readyz.timeoutMs` is `unhealthy` with `Probe timed out after …ms`, and `/readyz` answers `503`. On FrontMCP's Node server `enabled: false` and the paths behave the same way. [Answering health checks](https://frontmcp.dev/reference/server/observability#answering-health-checks) has every option.

---

## Troubleshooting

### `/readyz` answers `404`

The server runs as a Vercel or Lambda function, where readiness checks are off, or as a Worker, where they're off until you set `health: { readyz: { enabled: true } }`, or `health.enabled` is `false`. Use `/healthz`.

### `/health` or `/healthz` answers `404`

`health.enabled` is `false`, which removes every health path, or `healthzPath` moved the liveness check: `/healthz` is gone then, and `/health` answers in its place.

### Health checks get `403` `{"error":"Forbidden","message":"Invalid Host header"}`

The check calls the server by a name that isn't allowed, like `127.0.0.1:3000` or a pod's IP. Add that name to `FRONTMCP_ALLOWED_HOSTS`, or send the public name in a `Host` header, as in [the Kubernetes probes](#pointing-kubernetes-at-the-probes).

### `/readyz` answers `503`, with no details

In production the answer leaves `probes` out. Set `health: { includeDetails: true }` to see which probe failed, or check that Redis is reachable.

### `remote:…` is `degraded` right after startup

The server has a [remote app](https://frontmcp.dev/reference/server/remote), and FrontMCP hasn't checked it yet. `degraded` counts as ready, so `/readyz` answers `200`, and the probe turns `healthy` after the first check if the remote answers. (In 1.8.6 it was `unhealthy` until the first check and `/readyz` answered `503` for about 30 seconds.)

### `Failed to wire /metrics endpoint Error: Cannot find module '@opentelemetry/sdk-trace-base'`

`metrics.enabled` with `@frontmcp/observability` installed but not its peer. Install `@opentelemetry/sdk-trace-base`. Without `@frontmcp/observability` it's `Cannot find module '@frontmcp/observability'` ([more](https://frontmcp.dev/reference/server/observability#cannot-find-module-frontmcpobservability-at-startup)).

### The health check passes, but older clients' `initialize` answers `500`

The server is in production without `MCP_SESSION_SECRET`. Health checks, and clients on MCP 2026-07-28, don't notice. Set the secret.

### Prometheus reports the target `down` with `401` or `403`

`401`: the scrape sends no `Authorization` header. `403`: its token isn't `FRONTMCP_METRICS_TOKEN`. Check the `authorization` block and the credentials file.
