The payment provider fails. Your checkout counter keeps rising. The dashboard looks healthy. The trace contains a checkout span, but the payment call sits in another trace. And somewhere in the exception data, a customer email has slipped into your observability platform.

The application may have passed its functional tests. Its operational evidence has failed.

For Java teams, a useful next step in observability is to make telemetry part of the software contract: define what an important use case must emit, then test the actual SDK output before shipping it.

Why this matters now

OpenTelemetry graduated within CNCF in May 2026. Its Java documentation lists traces, metrics and logs as stable. That is a strong foundation for long-term engineering investment. It does not guarantee that a particular application's instrumentation answers the right operational questions. CNCF graduation announcement and Java signal status.

Two September updates make that gap especially relevant. On September 25, the project explained how instrumentation implementations can lag behind semantic conventions, with emitted data depending on configuration and application behavior. A dependency being present is therefore weaker evidence than observing what it produces. Instrumentation ecosystem update.

On September 22, its Prometheus interoperability survey reported improved usability alongside continued mixed deployments. The analysis included 81 screened active users from 186 respondents and excluded observability vendor employees. It is useful evidence about that sample, rather than a census of enterprise adoption. Survey results and methodology.

My editorial conclusion for Java and Spring teams is practical: invest in telemetry that remains understandable across library upgrades and pipeline changes. Installing another exporter is easy. Preserving the meaning of an incident dashboard requires deliberate engineering.

Start with the question an operator must answer

For a checkout use case, three questions are enough to begin:

  • How many attempts completed, and how many failed?
  • Did the payment operation belong to the same request trace?
  • Can we investigate the failure without collecting payment credentials or customer payloads?

Those questions become a small contract. Here is the contract used by the runnable example in this article:

Concern Contract Regression it detects
Business span One ended checkout.place span per valid attempt Missing or unfinished instrumentation
Failure ERROR status, failure outcome, bounded error classification Failed payments presented as healthy
Context Payment child span has the checkout span as parent An orphaned operation in another trace
Attempt metric One increment for success or runtime failure Failed attempts disappearing from the denominator
Dimensions WEB/MOBILE × success/failure Per-order metric streams
Payload handling No order ID or raw exception payload in this custom telemetry Sensitive content exported accidentally

A checkout use case produces a bounded attempt metric and a trace containing request, checkout and payment spans.

A use-case contract connects an operational question to concrete telemetry and an executable assertion.

The homann.* names below are application-specific attributes and a custom business metric. They are not OpenTelemetry HTTP semantic conventions. Keep standard HTTP and database instrumentation aligned with their published conventions; add domain telemetry where it supplies a different operational meaning.

A small Java implementation

The example uses Java 21 and OpenTelemetry Java SDK 1.66.0. It models a synchronous checkout application service with a payment port. There is no network request or database transaction in the test fixture.

package com.homannsoftware;

import io.opentelemetry.api.OpenTelemetry;
import io.opentelemetry.api.common.Attributes;
import io.opentelemetry.api.metrics.LongCounter;
import io.opentelemetry.api.trace.Span;
import io.opentelemetry.api.trace.StatusCode;
import io.opentelemetry.api.trace.Tracer;
import io.opentelemetry.context.Scope;
import java.util.Objects;

public final class Checkout {
    public enum Channel { WEB, MOBILE }
    public interface PaymentGateway { void charge(String orderId); }
    private final Tracer tracer;
    private final LongCounter attempts;
    private final PaymentGateway gateway;

    public Checkout(OpenTelemetry otel, PaymentGateway gateway) {
        this.tracer = otel.getTracer("com.homannsoftware.checkout", "1.0.0");
        this.attempts = otel.meterBuilder("com.homannsoftware.checkout").setInstrumentationVersion("1.0.0").build()
                .counterBuilder("homann.checkout.attempts")
                .setDescription("Completed checkout attempts, including failures")
                .setUnit("{attempt}").build();
        this.gateway = Objects.requireNonNull(gateway);
    }

    public void place(String orderId, Channel channel) {
        Objects.requireNonNull(orderId);
        Objects.requireNonNull(channel);
        Span span = tracer.spanBuilder("checkout.place").startSpan();
        span.setAttribute("homann.checkout.channel", channel.name());
        String outcome = "failure";
        try (Scope ignored = span.makeCurrent()) {
            gateway.charge(orderId);
            outcome = "success";
        } catch (RuntimeException ex) {
            // Raw exception messages may contain customer or payment data.
            span.setStatus(StatusCode.ERROR);
            span.setAttribute("error.type", "payment_failure");
            throw ex;
        } finally {
            span.setAttribute("homann.checkout.outcome", outcome);
            attempts.add(1, Attributes.builder()
                    .put("homann.checkout.channel", channel.name())
                    .put("homann.checkout.outcome", outcome).build());
            span.end();
        }
    }
}

The span name stays constant. The channel comes from an enum. The outcome has two values. The order identifier is passed to the payment port but never added to a span or counter.

On a runtime failure, the service preserves the original exception for its caller while emitting the bounded classification payment_failure. This intentionally trades detailed exception diagnostics for a predictable, payload-free signal. In a larger service, define a bounded failure taxonomy that distinguishes the incidents your operators need to handle.

This example does not call recordException(ex): exception messages and stack traces can include sensitive content. That choice applies to this custom instrumentation only. A Java agent, framework instrumentation, or logging bridge can still capture additional data, so inspect those outputs separately. OpenTelemetry's own guidance recommends minimizing collection and reviewing library instrumentation. Sensitive-data guidance.

makeCurrent() is equally significant. Creating a span does not make it the current parent for downstream instrumentation. The scope establishes that relationship for this synchronous call, then restores the previous context on both success and failure. An executor or asynchronous boundary needs its own propagation design and tests.

Test the emitted data, not calls to a mock

The companion Maven project uses InMemorySpanExporter, InMemoryMetricReader and a real SDK. It installs a synchronous span processor so completed spans are available immediately; the tests collect metrics explicitly. No backend or Docker container is required.

The fixture owns its SDK and closes it after each test. It does not register a global SDK. That keeps tests isolated and avoids depending on whichever test happened to initialize a singleton first.

This excerpt is taken directly from the executable test fixture:

private final InMemorySpanExporter exporter = InMemorySpanExporter.create();
    private final InMemoryMetricReader reader = InMemoryMetricReader.create();
    private final SdkTracerProvider traces = SdkTracerProvider.builder()
            .addSpanProcessor(SimpleSpanProcessor.create(exporter)).build();
    private final SdkMeterProvider metrics = SdkMeterProvider.builder()
            .registerMetricReader(reader)
            .registerView(InstrumentSelector.builder().setName("homann.checkout.attempts").build(),
                    View.builder().setAttributeFilter(Set.of(CHANNEL, OUTCOME))
                            .setCardinalityLimit(8).build()).build();
    private final OpenTelemetrySdk otel = OpenTelemetrySdk.builder()
            .setTracerProvider(traces).setMeterProvider(metrics).build();

    @AfterEach void close() { otel.close(); }

The metric view keeps only the two approved dimension keys and sets a deliberately small cardinality limit for the tests. OpenTelemetry Java supports both attribute filtering and cardinality configuration through views. Java SDK view configuration.

An allowlist of keys does not constrain their values. If someone replaces WEB with an order ID while keeping the same key, the view cannot infer that the value is wrong. That is why bounded domain types and output tests belong alongside SDK configuration.

One of the seven tests exercises 10,000 orders across all four channel/outcome combinations:

@Test void tenThousandOrderIdsStillProduceOnlyFourMetricPoints() {
        var ok = new Checkout(otel, id -> {});
        var fail = new Checkout(otel, id -> { throw new IllegalStateException("declined"); });
        for (int i = 0; i < 10_000; i++) {
            var channel = i % 2 == 0 ? Checkout.Channel.WEB : Checkout.Channel.MOBILE;
            if (i % 4 < 2) ok.place("order-" + i, channel);
            else assertThrows(IllegalStateException.class,
                    () -> fail.place("private-order", channel));
        }
        var points = points();
        assertEquals(4, points.size());
        assertEquals(10_000, points.stream().mapToLong(LongPointData::getValue).sum());
        assertTrue(points.stream().allMatch(p -> p.getValue() == 2_500));
        assertTrue(points.stream().allMatch(p -> p.getAttributes().size() == 2));
        assertEquals(10_000, exporter.getFinishedSpanItems().size());
    }

The full fixture also verifies span completion, metric unit and scope, original exception identity, exact custom error attributes, absence of exception events, parent-child trace relationships, context restoration after failure, and filtering of an unexpected order.id metric attribute.

The test pipeline runs a use case through a real SDK, collects spans and metrics in memory, and evaluates contract assertions.

The assertion boundary is the data leaving the SDK. A backend integration test supplies a second, separate boundary.

Cardinality protection can hide part of your failure rate

The example's two channels and two outcomes allow four attribute combinations in one SDK instance. Adding an order ID could turn that into one combination per order. Resource attributes, multiple instances, backend translation and retention add other dimensions to the production picture; four SDK points are not a promise of four backend time series globally.

The SDK's cardinality limit protects aggregation memory, but it changes the available breakdown when overflow occurs. Measurements that cannot retain their original attribute set enter an overflow aggregation marked otel.metric.overflow=true. For synchronous counters, the specification requires each measurement to be counted exactly once. The original dimensions are absent from the overflow point. Metrics SDK specification.

Our seventh test deliberately supplies 100 distinct channel values to the same counter. With the fixture's configured limit of eight, the tested Java implementation reports a seven-combination normal capacity and reserves an overflow point. The total remains 100; filtering for outcome=success returns less than 100 because that attribute is missing from overflow. We test this behavior rather than assuming how the implementation interprets its limit.

Four bounded metric combinations remain interpretable; unbounded values create an overflow point that retains counts but loses outcome labels.

Illustrative arithmetic, not a production benchmark: an intact total can coexist with an incomplete breakdown.

OpenTelemetry's August practical guide describes the same operational risk for dashboards and SLOs. Watch overflow signals and investigate unexpected dimension growth before increasing a limit. Cardinality guide.

Reproduce the evidence

Use the complete examples directory in the downloadable source package. It includes the Maven dependencies, source and all seven tests. With JDK 21 and Maven 3.9.11:

mvn -B clean test
python verify_mutations.py

On Windows, pass mvn.cmd to the Python runner if needed. Initial dependency resolution requires Maven Central. Python 3 is needed only for the optional mutation check.

Verified on October 1, 2026 with Eclipse Temurin 21.0.11+10 and OpenTelemetry SDK 1.66.0:

Tests run: 7, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS

missing-error-status REJECTED
missing-current-context REJECTED
exception-payload-leak REJECTED
Restored baseline exit code: 0

The mutation runner makes three controlled changes: remove the error status, remove checkout's current span context, and copy a raw exception message into the span. The unchanged contract tests reject all three. It restores the source in finally and reruns the complete suite. Use a disposable copy and avoid concurrent edits; forcibly terminating the process can interrupt restoration.

These results establish the behavior of this synchronous example and its pinned SDK. They do not establish collector delivery, Prometheus naming, asynchronous propagation, Java-agent coverage, sampling behavior, database instrumentation, backend retention, alert correctness, performance, or privacy across the whole application.

Fit the contract into a Spring system

Keep domain rules independent of the telemetry SDK. Instrument the application use case or an application-service decorator; keep infrastructure exporters and resource configuration at the composition boundary. The payment port remains an ordinary interface.

In production, reuse the application's configured OpenTelemetry instance and provider lifecycle. Do not copy the in-memory test bootstrap into a Spring application and accidentally create a second SDK beside an agent or starter. Choose which instrumentation owns each operation and check the resulting output for duplication.

Start with one important workflow, including its failure paths. Let developers and operators agree on the question, attribute vocabulary and required evidence. Add a contract test when the use case changes, review the dashboard against real incidents, and refine the contract as feedback arrives. Requirements, implementation, telemetry and operational experience inform each other throughout delivery.

For a release, use two levels of evidence: local contract tests for the service's emitted signals, and a representative integration check through the actual collector and backend. The second check should verify the translated metric name, labels, resource identity, query and alert behavior that people rely on during an incident.

A dashboard is a dependency of your incident response. Give its input the same care you give an API contract.

Researched and tested on October 1, 2026, Europe/Berlin. Topic selection is an editorial judgment for Homann Software's Java, Spring and architecture focus. Graphics are original explanatory assets; the featured illustration is AI-generated. All test evidence refers to the supplied example.