Reference

Benchmarks

Every run loads the same randomized data into two writers and a reader on MinIO, next to Tempo, then sends both the same randomized TraceQL queries. Each answer is checked for correctness and timed. The newest run is shown by default.

Runhistoricalvs Tempo on S3 2.10.8s3Full report ↗
Failed· 1 mismatch case(s)2026-10-08 04:24Z · 0b54fe965a · simd-perf · 30 min · OrbStack · aarch64 · 12 CPUs · 8.8 GiB
3,461
queries compared
1
mismatches
368 non-match total
5.88 ms
p50 query
410 ms Tempo
51.8 ms
p99 query
797 ms Tempo
0.02×
p50 latency ratio
per query, impl / oracle
0.03
mean CPU cores
0.15 Tempo
Latency by percentile
Plural TracesTempo on S3 2.10.8
Latency of every query both sides answered. Each query runs against both systems with the same data, so the percentiles compare like with like.
By query family
Query familyCasesp99 latencyp50 ratioNon-match
traceql.spanset1,195
62.2 ms
774 ms
0.02×171
trace.by_id608
23.1 ms
1.1 s
0.01×0
traceql.structural597
60.0 ms
762 ms
0.02×41
traceql.aggregate372
46.7 ms
1.0 s
0.02×16
tags.names156
22.3 ms
442 ms
0.02×99
tags.values153
26.7 ms
155 ms
0.04×1
traceql.by143
65.8 ms
672 ms
0.02×22
traceql.pipeline136
47.1 ms
635 ms
0.02×6
traceql.select101
30.2 ms
712 ms
0.01×12
Bars share one linear scale across families. The ratio is the median of per-query implementation / oracle latency, so values below 1× mean Plural answered faster.
Correctness
  • match3,093
  • inconclusive367
  • mismatch1
Inconclusive results are differences the API contract allows, such as both sides truncating at a limit.
Resources
CPU, mean cores
Plural
0.03
Tempo
0.15
Memory, p95 working set
Plural
125 MiB
Tempo
492 MiB
Sampled from the Docker Engine API every two seconds over the run. Object storage is excluded from both sides.
History · historical · Tempo on S3 2.10.8
Plural TracesTempo on S3 2.10.8
Only one run has been recorded on this scenario and oracle so far. Later runs will chart here.

How runs are measured

The differential fuzzer in tests/regression/harness/fuzz runs both systems on one Docker host, each pinned to its own CPU set. A run fails on any mismatch, implementation error or timeout, unstable implementation answer, or rejected write. The recent scenario queries data still being ingested; historical queries older data that has been flushed to object storage, with artificial latency added to MinIO.

To record a new run from the repository root:

bash
PYTHONPATH=tests/regression mise exec -- python -m harness.fuzz.bench --products traces --duration 30m

Results land in documentation/benchmarks/fuzz and appear here on the next docs build.