Token and cost tracking
Latitude aggregates token usage and estimated cost from each trace’s LLM spans. Use these fields to understand expensive interactions, compare models, and investigate behavior that drives high usage.What Latitude tracks
When the provider or framework instrumentation reports usage, Latitude shows:- input tokens
- output tokens
- total tokens
- estimated cost
- models used
- providers used
Find expensive traces
Filter by token usage or cost to find:- high total cost
- high input or output tokens
- expensive traces for one model or provider
- expensive traces within a tag, user, session, or metadata cohort
Where usage appears
Token and cost data appears in:- the Traces table
- trace detail views
- session aggregates
- filters
- search result rows
Where cost comes from
Latitude resolves each LLM span’s cost in this order:- Your own cost. The span carries cost attributes and
latitude.cost.sourceis set to"user". - Provider-reported cost. The instrumentation sent cost attributes without that marker.
- Estimate. Latitude prices the span’s token counts from public model pricing.
- Unpriced. Tokens were reported but no pricing matched the provider and model, so the cost stays at zero and the span is flagged as unpriced.
Cost attributes
All costs are in USD. Each value can be sent as a double, an int, or a numeric string such as
"0.0123".
Reporting your own cost
If you know what a call actually cost you, for example because of a negotiated rate, a fine-tuned model, or a self-hosted deployment, setlatitude.cost.source to "user" on the span along with any of the cost attributes above. Latitude then stores your figure exactly as sent and labels it as user reported. It never fills in a missing part from public pricing:
- Total only: the total is stored as sent, and the input and output costs stay at zero.
- Input and/or output only: the total is their sum. A side you leave out counts as zero.
- Both: the sides and the total are stored as sent.
0 is honoured, including a 0 total, so you can mark a call as free.
Any other value of latitude.cost.source, or no marker at all, means the cost is treated as provider reported. In that case a lone 0 total is ignored and the span is estimated, and a total sent without sides gets its input and output costs estimated from public pricing where it is available. If the marker is set but no valid cost attribute is present, the span is estimated as usual.
gen_ai.operation.name set to chat, text_completion, generate_content, embeddings or rerank/reranker (or the OpenInference, OpenLLMetry or Vercel AI SDK equivalent). Trace and session totals and the Cost page only count cost on those operations, so a cost on any other span is stored on that span but not counted. This applies to the telemetry SDKs’ set_llm_cost (Python) and setLlmCost (TypeScript) too; they log a warning, once per process, when called on a span that is not an LLM call.
Validation
Every cost attribute must be a finite number greater than or equal to zero. Negative values,NaN, infinities, and strings that are not numbers are ignored as if they had not been sent, and the span is still ingested. Latitude then uses the next attribute in the table that holds a valid value, or falls back to the estimate.
Related
- Percentile cohorts: Compare cost and token usage against similar tagged traces
- Traces: Trace-level usage and cost
- Sessions: Session-level aggregation
- Filters: Filter by cost and token usage
- Search: Discover costly behaviors