Traces

Architecture

A Tempo-compatible trace database on SlateDB. It takes OTLP, Zipkin and Jaeger, answers TraceQL and trace-by-ID lookups, and runs as two StatefulSets and a bucket.

Fig. 1 · Plural Traces vs Grafana Tempo
Query load
Plural Traces
two StatefulSets on SlateDB
AgentsOTel Collector · AlloyGrafanaquerieswriter1 shard / pod×2readerreads every shard×4Kubernetes LeasesSlateDBone per shardObject storage
2
component kinds
6
pods
0
replicated disks
0
hash rings

More writers means more shards from the next hour on. Nothing is copied or rebalanced, and readers scale on their own.

Grafana Tempo
microservices mode
AgentsOTel Collector · AlloyGrafanaqueriesdistributor×4ingesterRF=3 · live traces + WAL×9query-frontend×2querier×9metrics-gen×2memcached×3 pools×4memberlist rings: distributor · ingester · compactor · generatorcompactor×2Object storageParquet blocks
7
component kinds
32
pods
9 PVCs
replicated disks
4
hash rings

Ingesters hold live traces in memory, replicated three ways, until blocks are cut. Compactors shard work through their own ring; queriers fan searches across ingesters and backend blocks.

Pod counts are illustrative sizing for comparison, derived from vendor capacity guides and our provisional writer envelope. They are not benchmark results.

Write path

  1. A collector sends spans to any writer: OTLP over HTTP or gRPC, Zipkin v2 JSON, or Jaeger gRPC.
  2. Spans are grouped by trace ID, and each trace is hashed to a storage shard with the routing epoch in effect at its earliest span. Traces for other shards are forwarded to their owner over internal gRPC.
  3. The owner buffers traces in memory and acknowledges at the configured durability.
  4. On flush, traces are packed into pages of up to 1,024 traces (about 1 MiB). Page metadata, the compressed payload, trace heads and typed attribute postings are committed to SlateDB in one atomic write.
  5. Spans that arrive after their trace was flushed become a continuation page; the trace head records how many pages to read.

Read path

  1. Grafana sends a trace ID or a TraceQL search to any reader. Readers open every shard read-only.
  2. By ID: the trace head points at the trace's first page and page count, so a lookup reads exactly those pages and assembles one trace.
  3. TraceQL: equality predicates on typed attributes are resolved against postings to find candidate pages. Searches with no usable predicate prune pages by the per-trace time summaries in page metadata.
  4. Candidates are decoded and the complete TraceQL expression is evaluated, so the index never changes query semantics.
  5. Results are merged across shards and cut at the requested limit.

Storage layout

Each storage shard is one SlateDB database. Keys are scoped by tenant namespace and time segment:

Key layout
subsystem
0x05
version
0x06
namespace
bytes · 0x00
segment
i64 BE
record
u8
SlateDB segment ▲

Every key starts with this prefix, so each tenant namespace and segment is one contiguous key range. Retention drops whole ranges, and queries scan only the ranges in their time window. The record type comes next:

0x02Page metadataPer-trace time summaries for pruning
0x03Page payloadUp to 1,024 compressed OTLP traces
0x04Trace headTrace ID → first page and page count
0x05Attribute postingTyped attribute term → trace indexes
0x06Trace continuationLater pages of a long trace
0xffDiscovery catalogTag names and typed values
Fig. 2 · Every Traces key begins with the same prefix. Segments default to one hour.

Attribute posting keys include the scope (resource or span) and the value type, so the integer 7 and the string "7" never collide.

Compared with Tempo

ConcernTempoPlural Traces
Recent dataLive traces in ingester memory, 3× replicated, plus a WALWriter buffer, then SlateDB on object storage
Write fan-out3× to ingesters via the hash ring1× to the shard owner
Block formatParquet blocks, compacted through their own ringPages in SlateDB, compacted per shard
SearchQueriers fan out to ingesters and backend blocksReaders use typed attribute postings
Query tiersFrontend, queriers, compactors, cachesReaders

Compatibility

Against Tempo's TraceQL example corpus, 300 of 463 queries match Tempo exactly. Most of the rest are TraceQL metrics. Supported: scoped and unscoped attributes, every intrinsic, structural operators (>, >>, ~ and their negated and union forms), and count, sum, avg, min, max, by() and select() pipelines. Not supported:

  • TraceQL metrics (rate(), *_over_time, compare()). The query-range route returns 501.
  • Span event and link attributes (event.x, link.x).
  • A pipeline as a structural operand.