Skip to main content

Token and cost tracking

Latitude aggregates token usage and estimated cost from each trace’s LLM spans. Use these fields to understand expensive interactions, compare models, and investigate behavior that drives high usage.

What Latitude tracks

When the provider or framework instrumentation reports usage, Latitude shows:
  • input tokens
  • output tokens
  • total tokens
  • estimated cost
  • models used
  • providers used
Trace rows show aggregate token and cost values across their spans. Session rows aggregate usage across their traces.

Find expensive traces

Filter by token usage or cost to find:
  • high total cost
  • high input or output tokens
  • expensive traces for one model or provider
  • expensive traces within a tag, user, session, or metadata cohort
You can combine these filters with Search. For example, search for agent loops between tools and filter to traces above a cost threshold.

Where usage appears

Token and cost data appears in:
  • the Traces table
  • trace detail views
  • session aggregates
  • filters
  • search result rows
Unless your telemetry reports its own cost, cost values are estimates based on the usage and model/provider information available in telemetry.

Where cost comes from

Latitude resolves each LLM span’s cost in this order:
  1. Your own cost. The span carries cost attributes and latitude.cost.source is set to "user".
  2. Provider-reported cost. The instrumentation sent cost attributes without that marker.
  3. Estimate. Latitude prices the span’s token counts from public model pricing.
  4. Unpriced. Tokens were reported but no pricing matched the provider and model, so the cost stays at zero and the span is flagged as unpriced.

Cost attributes

All costs are in USD. Each value can be sent as a double, an int, or a numeric string such as "0.0123".

Reporting your own cost

If you know what a call actually cost you, for example because of a negotiated rate, a fine-tuned model, or a self-hosted deployment, set latitude.cost.source to "user" on the span along with any of the cost attributes above. Latitude then stores your figure exactly as sent and labels it as user reported. It never fills in a missing part from public pricing:
  • Total only: the total is stored as sent, and the input and output costs stay at zero.
  • Input and/or output only: the total is their sum. A side you leave out counts as zero.
  • Both: the sides and the total are stored as sent.
An explicit 0 is honoured, including a 0 total, so you can mark a call as free. Any other value of latitude.cost.source, or no marker at all, means the cost is treated as provider reported. In that case a lone 0 total is ignored and the span is estimated, and a total sent without sides gets its input and output costs estimated from public pricing where it is available. If the marker is set but no valid cost attribute is present, the span is estimated as usual.
Reported cost, yours or the provider’s, counts as verified spend on the Cost page. Estimated cost does not. Report the cost on the LLM-call span itself, with gen_ai.operation.name set to chat, text_completion, generate_content, embeddings or rerank/reranker (or the OpenInference, OpenLLMetry or Vercel AI SDK equivalent). Trace and session totals and the Cost page only count cost on those operations, so a cost on any other span is stored on that span but not counted. This applies to the telemetry SDKs’ set_llm_cost (Python) and setLlmCost (TypeScript) too; they log a warning, once per process, when called on a span that is not an LLM call.

Validation

Every cost attribute must be a finite number greater than or equal to zero. Negative values, NaN, infinities, and strings that are not numbers are ignored as if they had not been sent, and the span is still ingested. Latitude then uses the next attribute in the table that holds a valid value, or falls back to the estimate.
  • Percentile cohorts: Compare cost and token usage against similar tagged traces
  • Traces: Trace-level usage and cost
  • Sessions: Session-level aggregation
  • Filters: Filter by cost and token usage
  • Search: Discover costly behaviors