VALE Observability Metrics

This concept was last updated March 19, 2026

If you're running a system that people depend on, you need to observe how well it's working. Throughput is a popular measurement for this purpose, but it's unfortunately a poor metric.

Why? Because throughput is a noisy proxy for (Availability * Latency) when Volume > 0 and Errors ≈ 0, where:

  • Volume is the count of operations performed by a system

    "How much load is on the system?"

  • Availability is the gauge of the percent of traffic reaching the system

    "How much capacity is available?"

  • Latency is the distribution of operations' latency

    "How smoothly is the system operating?"

  • Errors is the count of failed operations

    "Is anything going really wrong?"

These are the VALE[1] metrics, and they form the core of observability (O11y).

To effectively gather and use VALE metrics, you'll need to understand three closely related concepts: SLIs, SLOs, and SLAs.

  1. I adapted VALE from the VALET framework, which adds a fifth metric: Tickets (manual interventions). VALE drops the T to focus on metrics that can be collected even without mature operational tooling. ↩

Service-Level Indicators (SLIs)

An SLI is a specific quality measured directly from a system. SLIs answer concrete questions about what's happening:

  • "How many inputs have been processed?"

  • "How fast was each input processed?"

  • "How often does the system reject an input?"

SLIs are the raw signal that every other business metric is built on top of, and they come in many shapes.

Most SLI metrics fall into two categories: Core metrics that are measured directly, and derived metrics that are calculated from other metrics.

Core Metrics

Counts measure continuous sums of something—the number of requests received, errors observed, or products processed by a system.

A count metric visualized over time

Gauges measure discrete values of something—the duration of a request, the size of an input, or the temperature of a CPU.

A gauge metric visualized over time

Derived Metrics

Rates measure a count over time—requests per second, errors per day, products processed per month.

A rate metric visualized over time

Histograms measure a gauge over time, capturing the mean, median, minimum, maximum, and rate of observed values.

A histogram metric visualized over time

In distributed tracing platforms like Datadog, distributions are centrally-aggregated histograms that enable accurate percentile computations across multiple systems.

A distribution metric visualized over time

Service-Level Objectives (SLOs)

An SLO is an aspirational, internal goal that an SLI should meet. SLOs are targets that teams set for themselves:

  • "We want each processing stage to take less than 96ms"

  • "We expect fewer than 42 errors per hour."

  • "We expect less than 0.01% of valid inputs to be rejected."

SLOs are ambitious by design: If you're always meeting them, they're probably too loose.

Service-Level Agreements (SLAs)

An SLA is an external agreement with customers that systems' operators—and their SLIs—must meet:

  • "We guarantee every request is processed in under 420ms"

  • "We expect less than 0.1% of valid requests will fail."

  • "We promise the mean-time-to-resolve errors will always be less than 9 hours."

SLAs should always be looser than your SLOs: Violating an SLA often carries a contractual, or at least social, cost.

How They Fit Together

Think of it as a stack:

  1. SLIs are the measurements.

  2. SLOs are the internal targets for those measurements.

  3. SLAs are the external commitments backed by those targets.

If your SLIs breach an SLO, your team investigates. If your SLIs breach an SLA, your users are impacted and your team is accountable.

When your SLIs consistently meet your SLOs, the remaining margin is your error budget—the amount of unreliability you can tolerate before breaching your objectives. Error budgets turn reliability into a resource you spend deliberately, rather than a target you chase reactively.