Architecture
A Tempo-compatible trace database on SlateDB. It takes OTLP, Zipkin and Jaeger, answers TraceQL and trace-by-ID lookups, and runs as two StatefulSets and a bucket.
- 2
- component kinds
- 6
- pods
- 0
- replicated disks
- 0
- hash rings
More writers means more shards from the next hour on. Nothing is copied or rebalanced, and readers scale on their own.
- 7
- component kinds
- 32
- pods
- 9 PVCs
- replicated disks
- 4
- hash rings
Ingesters hold live traces in memory, replicated three ways, until blocks are cut. Compactors shard work through their own ring; queriers fan searches across ingesters and backend blocks.
Write path
- A collector sends spans to any writer: OTLP over HTTP or gRPC, Zipkin v2 JSON, or Jaeger gRPC.
- Spans are grouped by trace ID, and each trace is hashed to a storage shard with the routing epoch in effect at its earliest span. Traces for other shards are forwarded to their owner over internal gRPC.
- The owner buffers traces in memory and acknowledges at the configured durability.
- On flush, traces are packed into pages of up to 1,024 traces (about 1 MiB). Page metadata, the compressed payload, trace heads and typed attribute postings are committed to SlateDB in one atomic write.
- Spans that arrive after their trace was flushed become a continuation page; the trace head records how many pages to read.
Read path
- Grafana sends a trace ID or a TraceQL search to any reader. Readers open every shard read-only.
- By ID: the trace head points at the trace's first page and page count, so a lookup reads exactly those pages and assembles one trace.
- TraceQL: equality predicates on typed attributes are resolved against postings to find candidate pages. Searches with no usable predicate prune pages by the per-trace time summaries in page metadata.
- Candidates are decoded and the complete TraceQL expression is evaluated, so the index never changes query semantics.
- Results are merged across shards and cut at the requested limit.
Storage layout
Each storage shard is one SlateDB database. Keys are scoped by tenant namespace and time segment:
Every key starts with this prefix, so each tenant namespace and segment is one contiguous key range. Retention drops whole ranges, and queries scan only the ranges in their time window. The record type comes next:
Attribute posting keys include the scope (resource or span) and the value type, so the integer 7 and the string "7" never collide.
Compared with Tempo
| Concern | Tempo | Plural Traces |
|---|---|---|
| Recent data | Live traces in ingester memory, 3× replicated, plus a WAL | Writer buffer, then SlateDB on object storage |
| Write fan-out | 3× to ingesters via the hash ring | 1× to the shard owner |
| Block format | Parquet blocks, compacted through their own ring | Pages in SlateDB, compacted per shard |
| Search | Queriers fan out to ingesters and backend blocks | Readers use typed attribute postings |
| Query tiers | Frontend, queriers, compactors, caches | Readers |
Compatibility
Against Tempo's TraceQL example corpus, 300 of 463 queries match Tempo exactly. Most of the rest are TraceQL metrics. Supported: scoped and unscoped attributes, every intrinsic, structural operators (>, >>, ~ and their negated and union forms), and count, sum, avg, min, max, by() and select() pipelines. Not supported:
- TraceQL metrics (
rate(),*_over_time,compare()). The query-range route returns501. - Span event and link attributes (
event.x,link.x). - A pipeline as a structural operand.