Observability should be Cheap
Why we're building Plural Telemetry, and the principles every design decision is held to.
Rethinking Observability in an Agentic Era
Since I've operated significant cloud infrastructure, the most common painpoint were observability bills. The standard tradeoff was between:
- extreme operational toil managing extreme high write throughput telemetry stacks
- extreme Datadog bills to manage it for you
Almost everyone picked 2 because not only do they do a great job, but they provided a great product on top. We think that world has now changed. There are a few key things that have happened:
- Agents are going to be the primary consumers of observability data and they just need a programmable REST api to use. The real value is at the data layer.
- Observability shaped data is a perfect fit for the new breed of LSM tree backed, object storage native databases, making the prior operational toil much more tractable
We've invested and will continue to invest in both. Plural provides a full agentic solution for platform engineering, with observability being a core integration point. But also the Plural Telemetry project is meant to fully push the boundaries of leveraging Slatedb to bring an operationally simple observability database toolkit on top of vanilla K8s and object storage.
The goal is all you should need to get high quality, fast telemetry is 3 yaml blobs, an s3 bucket and a small AWS bill.
Key Design Principles
1. Object storage is the database
S3 provides infinite durable storage, and slatedb manages storage tiering into NVMe for warm caching. This allows a simple reader/writer architecture, and infinite scalability.
2. Fewer moving parts
The existing solutions all have operational warts, the LGTM stack has an seriously complex microservices architecture, need for external caching in memcached, and terrible tail latencies. VictoriaMetrics still requires manually managing durable storage, alongside its own sharding complexities.
Plural Telemetry is straightforward - you scale sharded writers to match ingest, and scale readers to match reads. Everything else synchronizeed from a single s3 bucket.
3. Compatible by default
Re-instrumenting is a non-starter. Each database speaks the protocols you already run: Prometheus remote write, Loki push, OTLP, Zipkin, Elasticsearch _bulk. Each answers PromQL, LogQL or TraceQL through the same HTTP APIs Grafana already uses. Switching should be a change of URL.
4. Kubernetes is the control plane
Distributed databases usually bring their own consensus system. Kubernetes already is one, and it's already running. Plural Telemetry coordinates through StatefulSets for stable identity, a ShardMap custom resource for routing, and Lease objects for shard ownership, so there's no ZooKeeper, etcd or gossip protocol to operate.
5. Scale without moving data
Rebalancing is where distributed stores get dangerous. Because telemetry is ordered by time, Plural Telemetry scales by changing where future data goes. A new routing epoch takes effect at the next hour boundary, new shards start empty, and nothing is copied. See epoch sharding.
6. Prove it
Claims about speed and compatibility are cheap. Every database is fuzzed against its reference implementation on identical data, and the latency, CPU and memory of both sides are recorded and published, including the cases where we lose. See each database's Benchmarks page.
7. Open, and sponsored by Plural
Plural Telemetry is Apache 2.0 and developed in the open by Plural. Plural runs fleets of Kubernetes clusters for its customers, and cheap, boring observability is a prerequisite for doing that well.
What we won't do
- Build a proprietary query language. If PromQL, LogQL and TraceQL can express it, that's what we implement.
- Require a managed service. Everything runs in your cluster, against your bucket.
- Hide the trade-offs. Writer shards aren't replicated in-cluster: failover moves a
Leaseand reopens the same data from object storage. Scale-down isn't supported yet. Both are documented rather than glossed over.
Get involved
Read the architecture guide, try the installation, and open issues or pull requests on GitHub.