Record your own AI agents
Everything above is about the people visiting your site. This section is about the agents you run yourself: what they ran, which models they called, how many tokens that used, and what it cost. It is included on every plan and shares the same account, key scopes and event meter as your web analytics.
A run is one unit of work, such as a code review or a support reply. Inside it, each model call, tool call or retrieval is a span with its own timing, status and, for model calls, token counts. Cost is worked out on the server from those token counts and the model's pricing.
Install the SDK
The agent telemetry SDK is available in the project's sdk/ source directory,
but is not currently published on npm. Build it from a repository checkout before using the
examples below. Website tracking and MCP access do not require this package.
The initialiser identifies the framework, inserts the browser tracker, and writes an SDK bootstrap if the project has AI dependencies. If it cannot identify the project it prints the snippet and exits non-zero rather than guessing, because a snippet written somewhere plausible and wrong shows up as missing data weeks later.
Report a run
Wrapping an Anthropic client records every call it makes without touching the call sites. Tool calls are one line each around the work they time.
import Anthropic from "@anthropic-ai/sdk";
import { StepMetrics, instrumentAnthropic } from "@stepmetrics/sdk";
const analytics = new StepMetrics({
endpoint: "https://app.stepmetrics.co",
apiKey: process.env.STEPMETRICS_API_KEY, // an ingest-scope key
siteId: "<SITE_ID>",
environment: "production",
agent: "code-review",
});
const run = analytics.startRun({ name: "code-review" });
const client = instrumentAnthropic(new Anthropic(), run);
const search = run.tool("search_code");
const hits = await searchCode(query);
search.end();
await client.messages.create({ model: "claude-sonnet-5", ... }); // recorded as an llm span
await run.end();
Any other provider, or a loop you wrote yourself, reports the same way: open a span with
run.llm({ model, provider }), then span.end({ inputTokens,
outputTokens }) with the numbers the provider returned, or span.fail(err).
Spans are batched and sent in the background, and a run flushes what is left when it ends.
Read it back
Agent telemetry is read over MCP and the API; it does not appear in the web dashboard. The
five MCP tools are get_agent_overview, get_agent_models,
get_agent_tools, get_agent_timeseries and
get_agent_failures. The same five answers are at
/v1/sites/<SITE_ID>/agent/overview, /models,
/tools, /timeseries and /failures, and every one
takes the same period as the web endpoints plus an optional
environment and run name filter.
Cost, and what happens without a price
Send tokens, not money. Cost is computed on the server from the provider's own token counts. A model with no pricing row is recorded as unknown cost rather than as zero, and the ingest response tells you which, so an undercount is visible while you can still fix it.
Retries and double-counting
Span keys are generated per run, so a replayed run regenerates the same keys and the server writes over the originals rather than appending a second copy. If your run is nondeterministic, pass your own run key to make that hold.
Ingest keys
Write-only, and their own scope rather than a weaker read key. A leaked ingest key can add telemetry and do nothing else: it cannot read a single row, mint a key, or delete anything.