Benchmarks
Every run loads the same randomized data into two writers and a reader on MinIO, next to Tempo, then sends both the same randomized TraceQL queries. Each answer is checked for correctness and timed. The newest run is shown by default.
| Query family | Cases | p99 latency | p50 ratio | Non-match |
|---|---|---|---|---|
| traceql.spanset | 1,195 | 62.2 ms 774 ms | 0.02× | 171 |
| trace.by_id | 608 | 23.1 ms 1.1 s | 0.01× | 0 |
| traceql.structural | 597 | 60.0 ms 762 ms | 0.02× | 41 |
| traceql.aggregate | 372 | 46.7 ms 1.0 s | 0.02× | 16 |
| tags.names | 156 | 22.3 ms 442 ms | 0.02× | 99 |
| tags.values | 153 | 26.7 ms 155 ms | 0.04× | 1 |
| traceql.by | 143 | 65.8 ms 672 ms | 0.02× | 22 |
| traceql.pipeline | 136 | 47.1 ms 635 ms | 0.02× | 6 |
| traceql.select | 101 | 30.2 ms 712 ms | 0.01× | 12 |
- match3,093
- inconclusive367
- mismatch1
How runs are measured
The differential fuzzer in tests/regression/harness/fuzz runs both systems on one Docker host, each pinned to its own CPU set. A run fails on any mismatch, implementation error or timeout, unstable implementation answer, or rejected write. The recent scenario queries data still being ingested; historical queries older data that has been flushed to object storage, with artificial latency added to MinIO.
To record a new run from the repository root:
PYTHONPATH=tests/regression mise exec -- python -m harness.fuzz.bench --products traces --duration 30mResults land in documentation/benchmarks/fuzz and appear here on the next docs build.