top of page

Ingest Everything, Understand Nothing: Why the Pipeline Is Now Your Most Important Cost Control

Sep 25
8 min read

For the past decade, “collect everything and decide later” has been the default observability strategy. It was understandable. Storage was becoming cheaper, telemetry standards were maturing and teams wanted to preserve every possible clue for the next incident.

That default is now broken.

A Gartner report published on 2 September 2026 describes how organisations have routinely ingested all operational data with the intention of analysing it later. Rising ingestion and storage costs have made that approach increasingly unworkable. The report also highlights an uncomfortable consequence: high volumes of poor-quality telemetry can reduce the value of generative AI and large language models that depend on it.

The market is changing rapidly. The cost of keeping everything is becoming more visible, while the operational value of much of that data remains uncertain. The era of “ingest now, understand eventually” is coming to an end.

But the obvious response, simply collecting less, can create an even greater risk.

The end of collect everything, decide later

Modern digital estates generate data at a scale that traditional collection assumptions cannot absorb indefinitely. Public cloud services, Kubernetes clusters, containers, serverless workloads and managed platform services all produce telemetry through different mechanisms, at different rates and with different operational meaning.

A short-lived Kubernetes pod may exist for seconds. A serverless function may create a burst of traces without leaving a persistent host footprint. A managed database may expose useful service metrics but restrict the underlying infrastructure context. High-cardinality labels can multiply metric volume, while application teams may emit verbose logs during every deployment.

The result is an unforgiving telemetry environment. Data volumes rise quickly, dependencies change constantly and the cost of retaining everything compounds across every downstream platform.

The Gartner finding is a stark reminder that this is not only a finance problem. Bad telemetry also weakens the systems now being asked to interpret it. Duplicated, inconsistent, poorly enriched data makes it harder for AIOps platforms, LLM-driven analysis and agentic investigation workflows to distinguish a meaningful signal from background noise.

More data does not automatically create more insight. Sometimes it creates a more expensive way to become uncertain.

Why indiscriminate reduction is not governance

When the invoice rises, the first instinct is often to cut volume. That instinct is understandable, but reduction without design discipline is not observability governance.

Filters are frequently written by whoever happens to be on shift. They may be undocumented, unreviewed and never revisited. A rule that appears harmless can quietly remove the one event, attribute or trace branch required to investigate the next major incident.

Sampling creates a similar problem. It may be statistically sound across a large population while discarding the exact rare event that matters most: the failed payment, the slow customer journey, the unusual error path or the anomalous edge case.

The monthly cost report may look excellent. The post-incident review may tell a very different story when the evidence is gone and cannot be reconstructed.

This risk is particularly serious when the same reduction logic is applied uniformly to logs, metrics and traces. These signals have different value densities and different investigative purposes:

  • Logs may require pattern clustering, sensitive-data detection and selective retention.

  • Metrics may benefit from aggregation, compaction and careful control of cardinality.

  • Traces may need high fidelity for selected services, transactions or customer journeys, even when broad sampling is appropriate elsewhere.

The danger of dropping data is rarely immediate. It appears weeks or months later, at the worst possible moment, when the missing telemetry is unrecoverable.

Cloud-native telemetry from Kubernetes, containers, serverless and managed services flowing into a qualification fabric

Pre-ingestion qualification is an architectural decision

The answer is not to preserve everything forever or to discard everything aggressively. The answer is to make the decision before ingestion, deliberately and with the same rigour applied to any other architectural choice.

That begins with purpose.

Before collecting a stream, teams should be able to explain which use cases it supports, who consumes it, which decisions depend on it and what retention obligations apply. Data with no defined purpose is not automatically future-proofing. It may simply be cost.

Signal classes should then be treated differently. High-value traces and security events may justify full fidelity and long retention. Metrics may be aggregated and compacted. Verbose debug logs may have a short searchable life, with longer-term replay-anything storage used only where the business case supports it.

Placement matters as much as selection. High-value data can be routed to full-fidelity analysis. Long-term archives can be sent to object storage. Lower-value volume can be placed into cheaper retention tiers. However, an archive is only useful if it can be recovered within the time required by the investigation. Replay must be designed and tested, not assumed.

Enrichment belongs in the pipeline as well. Adding service ownership, environment, geography, deployment and customer context while data is in flight makes telemetry more useful and more economical. Unenriched data is expensive twice: it costs money to store, and it costs engineers time to interpret.

Governance should also happen before storage. PII masking, sensitive-data detection, RBAC and versioned configuration should be built into the pipeline rather than bolted onto a destination platform after data has already travelled through the estate.

The pipeline itself requires observability. Its configuration needs review, versioning and rollback. Its throughput, failure rates, latency and queue depth need monitoring. A silent pipeline failure is an outage in the organisation’s ability to investigate outages.

Finally, quotas and budgets must become first-class controls. A runaway source should be caught by policy, not discovered when the invoice arrives.

The discipline behind every drop decision

Before removing a field, event, span or time series, we recommend asking three questions:

Which use case or decision depends on this data?
If we drop it, can we replay or recover it, and how quickly?
Who signed off the qualification rule, and when was it last reviewed?

These questions separate cost engineering from cost cutting.

They also create a practical control loop. A rule can be evaluated against current incidents, changing architecture, security requirements and actual query patterns. If a team cannot identify the owner or the review date, the rule is not governed, regardless of how sophisticated the filtering technology appears.

The pipeline is now an AI-enablement layer

Telemetry has new consumers. Humans still use dashboards and search, but models and agents increasingly inspect the same operational data.

That makes upstream quality more important. Poorly structured, duplicated or unenriched telemetry degrades LLM-driven analysis, agentic investigation and AIOps correlation. The resulting recommendations may be slower, less precise or based on a misleading concentration of low-value events.

Quality in, intelligence out.

This is becoming more urgent as agent workloads generate their own telemetry. Each prompt, tool call, handoff, retrieval step and model response may be traced. Without qualification, agentic systems could create the next observability cost cliff: huge volumes of technically interesting data with unclear retention, privacy and operational value.

The right question is not whether every interaction should be recorded indefinitely. It is which agent behaviours must be investigated, evaluated, audited or replayed, and which can be sampled, summarised or retained in a lower-cost tier.

Clean enriched telemetry feeding an AI reasoning core while noisy data creates interference

Use the pipeline as a decision framework, not just a product category

Our Telemetry Pipeline Periodic Table provides a practical framework for this work. It covers 42 capabilities across 18 vendors, with logs, metrics and traces tracked separately across ingestion, processing and forwarding.

The relevant capabilities include Filter & Drop, Dedup & Aggregate, Pattern Clustering, Logs to Metrics, Quotas & Budgets, Enrichment, PII Masking, Sensitive Detect, Versioning, Archive & Replay, Routing, Buffering and Pipeline Observability. It also includes AI Parsing, AI Optimisation, MCP and Agent API capabilities, which are becoming increasingly relevant as telemetry pipelines support automated analysis and rule authoring.

The comparison also separates vendors that meter pipeline throughput from those that do not. Edge Delta made throughput free at any scale in April 2026, while Bindplane increased its free tier in July 2026. Those changes demonstrate why licensing assumptions should be validated rather than copied from an old business case.

Hosting matters too. The table distinguishes SaaS, BYOC/hybrid and on-premise models. Its forwarding tiles measure vendor-neutral delivery, showing whether a pipeline can route logs, metrics and traces to third-party observability platforms, SIEMs, object storage or OTLP endpoints, rather than only to the vendor’s own backend.

We called this. The warning is that most teams still have not operationalised it.

In our 2023 observability predictions post, we identified observability pipelines as our number one theme for the year and noted that Gartner had not mentioned them at all, even as Datadog’s early move suggested the wider market would eventually follow. We returned to that argument in our later OpenTelemetry post, where we framed telemetry pipelines as the control point for sampling, filtering, tiered storage and data placement.

That prediction has since been borne out. In the August 2024 Gartner Magic Quadrant for Observability Platforms, telemetry pipelines appeared as a crucial feature for aggregating, processing and routing data across systems, reducing lock-in and supporting data federation in large enterprises. Since then, the category has expanded rapidly. What was once a niche adjacent capability is now a competitive market with dedicated pipeline products, sovereign and BYOC deployment models, AI-assisted rule authoring and agentic optimisation. Our own Telemetry Pipeline Periodic Table tracking 42 capabilities across 17 vendors is, in itself, evidence of how quickly the market moved.

But this should not be read as a victory lap. In our view, being right about pipelines is only useful if organisations act on it. The prediction was correct. The missing discipline is pre-ingestion qualification. Without that, a pipeline is just another route into uncontrolled cost, weak governance and noisy AI inputs.

Our Observability Periodic Table is the companion resource for assessing destination platforms. It covers 42 capabilities across 40 observability vendors, including data pipelines, cloud monitoring, Kubernetes, AI observability, automation, compliance and cost management.

Neither table is a universal ranking or a substitute for use-case validation. They are evidence-based comparison and decision-support tools. The starting point should remain the business outcome, priority use cases, payers, users and change-management capacity. Vendor claims should then be tested against the capabilities the organisation actually needs.

Own the telemetry; negotiate from a position of strength

Organisations should own their telemetry and their data. That may mean retaining it in their own cloud, their own datacentre or a suitable third-party cloud, while using vendor interfaces for advanced analysis and response.

The vendor platform should not necessarily be the only home for every event.

Pipelines make that separation possible. They provide portability between observability platforms, SIEMs, archives and AI systems. Portability is what makes the cost conversation negotiable. It allows teams to change destinations, retain valuable data independently and avoid treating every increase in platform pricing as an unavoidable infrastructure tax.

The objective is not the smallest possible telemetry bill. It is the smallest bill that still supports the operational, security, compliance and AI use cases the business depends on.

That is doing more with less, not simply doing less.

Governed telemetry pipeline with privacy controls, versioning, budgets and archive replay

A controlled asset, not a growing invoice

Pre-ingestion qualification is now a core part of observability strategy. It determines what the organisation can know during an incident, what its AI systems can understand and how much it pays to preserve both.

In our view, the strongest teams will treat telemetry as a governed product: purpose-led, enriched, portable, recoverable and continuously reviewed. They will monitor the pipeline as carefully as the services it supports, and they will prioritise the evidence required for decisions rather than rewarding volume for its own sake.

We help organisations turn complex telemetry estates into controlled, governed assets: from modern cloud and Kubernetes environments to legacy platforms, managed services and agentic workloads. Speak to our team about optimising your pipeline design, validating vendor claims and building a qualification model that supports both operational confidence and sustainable cost.

Observe. Automate. Operate with confidence.

 
 
 

Comments


bottom of page