Virtual threads make it practical to serve many blocking requests at once. Your inventory service, payment provider and database still have finite capacity. When a dependency slows down, an application that can start more work may also become very good at overwhelming it.
Spring Framework 7 gives teams useful building blocks for this boundary: native retries and method concurrency limits. The interesting engineering question is where to place them, which failures deserve another attempt, and what happens when the application is already full.
This article builds a small inventory lookup around those decisions. The accompanying Maven project runs against Spring Framework 7.0.9 and Java 21, with real Spring proxies and virtual threads. Fourteen integration tests exercise failure and admission behavior. Three deliberately introduced regressions are detected by the tests.
Sources checked on October 7, 2026. The tests use a controlled inventory adapter; they do not call a live service or measure production throughput.
Why this is the Spring topic I would prioritize now
The current Java and Spring landscape offers several attractive directions. Java 27 became available on September 15. Spring's documentation lists Boot 4.1.1 as stable and 4.2.0-M2 as a preview. Project Valhalla has value objects integrated for the JDK 28 preview cycle. These developments deserve attention, with a clear distinction between released software and early access. See the Java 27 announcement, Spring version listings, and OpenJDK's Valhalla project update.
For a team delivering business software today, I would put dependency protection ahead of experimenting with another language feature. Homann Software has already explored Java 27 structured concurrency and verifiable Spring Modulith boundaries. Resilience adds the operational question: can those boundaries remain useful when callers multiply and dependencies fail?
This is an editorial choice based on practical relevance, rather than a popularity ranking. Spring's resilience support arrived with Framework 7.0; it is not a feature newly released in October. The timely opportunity is to turn it into an explicit, tested service contract. Spring's September 21 release-process update also reinforces the need to track the exact versions and patch cadence used in production.
Start with the inventory contract
Assume a checkout flow needs a current inventory reading. The example makes four decisions:
- A temporary upstream failure can trigger two retries after the first attempt.
- Invalid input fails immediately.
- At most three synchronous adapter calls can be active through one gateway bean at a time.
- When all three slots are occupied, another caller receives a rejection instead of waiting at that gateway.
The numbers are deliberately small so the tests can expose the behavior. They are not recommended production settings. A real limit needs evidence about provider capacity, connection pools, request deadlines, replica count and other consumers of the same dependency.
The retry coordinator calls a separate Spring bean for every attempt. A slot covers the synchronous adapter call. Backoff happens after that call has returned or thrown.
Virtual threads change the cost of waiting for blocking I/O; they do not create additional database connections or remote-service capacity. OpenJDK explicitly discusses limiting access to scarce services independently of pooling virtual threads in JEP 444. That distinction is the motivation for the boundary here.
Put admission control around one adapter attempt
The gateway delegates to an InventoryPort, a small interface whose read(String sku) method returns an inventory result. The test implementations can block, succeed or throw. In production, an implementation would call the inventory transport and classify its failures.
package com.homannsoftware.resilience;
import org.springframework.resilience.annotation.ConcurrencyLimit;
import static org.springframework.resilience.annotation.ConcurrencyLimit.ThrottlePolicy.REJECT;
public class InventoryGateway {
private final InventoryPort port;
public InventoryGateway(InventoryPort port) {
this.port = port;
}
@ConcurrencyLimit(limit = 3, policy = REJECT)
public String read(String sku) {
return port.read(sku);
}
}
REJECT matters. In the pinned version, the default policy is BLOCK; REJECT throws InvocationRejectedException when the limit is reached. The policy attribute was added in Framework 7.0.3, so examples using it require at least that version. These details are documented in the versioned concurrency-limit API.
The test holds three calls inside the adapter, then makes twenty additional calls on virtual threads. All twenty are rejected before entering the adapter. The recorded peak is three active adapter calls. Once the held calls finish, another call succeeds. Separate tests check slot release after exceptions and after a cooperative interruption.
A rejected call still needs a product decision at the caller. Checkout might show a temporary inability to confirm stock; another use case might use a clearly marked cached reading. Translate that outcome at the application boundary. Returning invented availability to preserve a happy path would break the inventory contract.
This gateway is synchronous. If read starts background work and returns before that work ends, the slot no longer describes the lifetime of the resource operation. The same design should not be copied onto arbitrary asynchronous methods without testing their completion semantics.
Retry only a failure you understand
The coordinator uses Spring's native programmatic retry API. This makes the order of operations visible: retry invokes the proxied gateway; the gateway admits or rejects one attempt; the adapter performs the work.
package com.homannsoftware.resilience;
import java.time.Duration;
import org.springframework.core.retry.RetryPolicy;
import org.springframework.core.retry.RetryTemplate;
public class InventoryLookup {
private final InventoryGateway gateway;
private final RetryTemplate retry = new RetryTemplate(
RetryPolicy.builder()
.includes(TemporaryUpstreamFailure.class)
.maxRetries(2)
.delay(Duration.ofMillis(25))
.jitter(Duration.ofMillis(10))
.multiplier(2)
.maxDelay(Duration.ofMillis(100))
.build());
public InventoryLookup(InventoryGateway gateway) {
this.gateway = gateway;
}
public String find(String sku) {
return retry.invoke(() -> gateway.read(sku));
}
}
TemporaryUpstreamFailure is our own exception type. It expresses an adapter decision that the failure is transient and another read is appropriate. A real adapter should make that decision from the provider's documented behavior. It should not turn every HTTP error, authentication failure or malformed response into this exception.
The tests confirm three outcomes:
| Adapter behavior | Result | Adapter calls |
|---|---|---|
| Two temporary failures, then success | Inventory result returned | 3 |
| Temporary failure on every attempt | Last original failure propagated | 3 |
Invalid SKU reported as IllegalArgumentException |
Immediate failure | 1 |
We also count gateway invocations in a focused overload test. A local admission rejection causes one invocation of the retry callback. The retry policy does not include that exception. Retrying local overload immediately would simply ask the same full gateway to reconsider.
The configured backoff adds delay between eligible failures. The important structural property is that a failed adapter call has left the gateway before the delay begins. Waiting for a retry does not hold one of these three slots. Another caller may use the slot during that interval. The project tests admission behavior; it does not benchmark fairness or validate a statistical jitter distribution.
Spring's resilience reference describes both native annotations and RetryTemplate. The programmatic approach is useful here because retry ownership and the call to the admission boundary remain explicit.
Enable the proxy—and test the path that reaches it
The configuration is part of the example, not a step left to the reader:
package com.homannsoftware.resilience;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.resilience.annotation.EnableResilientMethods;
@Configuration(proxyBeanMethods = false)
@EnableResilientMethods
public class ResilienceConfig {
@Bean
InventoryPort inventoryPort() {
return sku -> "available:" + sku;
}
@Bean
InventoryGateway inventoryGateway(InventoryPort port) {
return new InventoryGateway(port);
}
@Bean
InventoryLookup inventoryLookup(InventoryGateway gateway) {
return new InventoryLookup(gateway);
}
}
@EnableResilientMethods enables processing of the method resilience annotations. The InventoryGateway injected into InventoryLookup is the Spring-managed proxy. Our test checks that with AopUtils.isAopProxy and calls through the actual container.
Framework 7 also has a native @Retryable annotation in org.springframework.resilience.annotation. Do not accidentally import the older Spring Retry project's annotation when migrating code. Our annotation test fixture uses the new package and observes the same three-attempt behavior for maxRetries=2.
Proxy routing creates a trap worth demonstrating. A method on a bean calls its own annotated read() method. That internal call never crosses the retry proxy, and the adapter fails after one invocation. With resilience processing disabled, a direct call to the annotated method also runs only once. The accompanying tests demonstrate both cases.
Keeping the coordinator and gateway in separate beans makes the intended route easier to inspect. Refactoring them into a single class with an internal call can silently remove the admission boundary unless the implementation changes accordingly. Spring explains self-invocation in its AOP proxy documentation.
Give retry ownership to one layer
A retry policy can look harmless in isolation. A caller making three attempts, calling another layer that makes three attempts, calling an adapter that also makes three attempts, can produce 27 leaf calls for one failed request.
Our test nests three real RetryTemplate instances around a persistently failing operation. Each permits two retries, and the leaf counter reaches exactly 27. This is a constructed failure scenario, not a measured production traffic pattern.
Worst-case call counts for the tested policies: three attempts at one layer; 3 × 3 × 3 leaf calls at three layers.
Choose the retry owner deliberately. Include SDK defaults, HTTP clients, service meshes, queue redelivery and job scheduling in that conversation. If several layers can repeat the operation, write down the resulting upper bound and the conditions under which each layer stops.
Reading inventory is a relatively simple example because repeating the read does not intentionally create a business side effect. Repeating a payment or shipment request needs a separate duplicate-handling contract. A timeout can occur after the receiver has committed work. Homann's durable workflow article explores that receiver-side problem; this inventory sample does not solve it.
A retry budget is not a transport deadline
The native retry API also supports a timeout. Its name can tempt readers to assume that it interrupts a slow operation.
We test that assumption directly. A RetryTemplate with a one-millisecond timeout starts a callback that waits on a latch. After the callback has entered, the test waits fifty milliseconds for the result. The result is still pending. When the latch is released, the original call returns successfully.
That experiment establishes a narrow but important fact for this synchronous API and version: the retry timeout does not preempt the callback already running. It should not be presented as a substitute for configuring the transport's own timeouts.
These controls govern different parts of a request. Their budgets must fit together, and cancellation must reach the actual adapter.
For a production inventory adapter, specify connection and response/read timeouts in the chosen HTTP client, account for connection acquisition where applicable, and test what cancellation does. Decide how much of the caller's remaining deadline another attempt can consume. The coordinator shown above intentionally demonstrates retry classification and admission; it does not implement propagation of an end-to-end deadline.
Neither the retry configuration nor the concurrency annotation is a circuit breaker. This sample also has no distributed quota or request-per-second limiter. If your contract needs those controls, evaluate them separately and test their interaction with this boundary.
A local limit needs a deployment calculation
Another test creates two independent Spring contexts, each with its own gateway and a limit of three. Six adapter calls can be active across those contexts. A seventh call is not globally counted against a shared limit; each gateway decides from its own local state.
This gives a useful deployment thought experiment. Four replicas with one such gateway each could admit up to twelve simultaneous calls through those gateways. Other applications may be calling the same provider as well. A limit that looks comfortable in one developer process can exceed the dependency's budget after scaling out.
Treat the annotation as local protection. Document how per-instance limits relate to the fleet and the provider. If the provider requires a hard shared ceiling, you need a coordinating mechanism at the appropriate boundary. Adding another annotation to each replica does not create that coordination.
Reproduce the evidence
The companion source project contains the full Maven configuration, application classes, test fixtures, mutation runner and captured outputs. It requires a JDK supporting Java 21 and Maven; dependencies are downloaded from Maven Central on the first run.
mvn -B clean test
The verification run used Eclipse Temurin 21.0.11+10, Maven 3.9.11, Spring Framework 7.0.9 and JUnit Jupiter 5.13.4. Surefire reported:
Tests run: 14, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS
To check whether the tests detect meaningful regressions, run:
python verify_mutations.py mvn.cmd
The runner temporarily removes the concurrency annotation, adds an extra retry, and broadens the retry classification to every runtime exception. Each mutation causes an assertion failure in the relevant test. It restores the original sources and runs the full suite again. The project README also gives a command to execute the demo, which prints available:SKU-42.
These checks establish behavior inside a real Spring context. They do not establish provider recovery rates, HTTP timeout handling, replica coordination, production latency, throughput or business-side-effect safety. The latch timeouts are test safeguards, not suggested service deadlines.
Turn the settings into architecture decisions
Before putting this pattern into a checkout service, I would record the dependency's permitted load, the local and fleet limits, the retry owner, retryable failure classifications and the user-visible overload outcome. Then I would connect each decision to a test and to production observations.
Track logical lookups separately from adapter attempts. Observe rejected admission, active adapter calls, retry exhaustion, transport timeouts and eventual user outcomes. A system can keep returning successful responses while consuming several times more provider calls per lookup; a success counter alone would miss that deterioration.
Spring Framework 7 makes the mechanics accessible. The valuable work is choosing a contract that preserves capacity and gives callers an honest outcome when the dependency cannot keep up. Start with one bounded adapter, make the failure paths reproducible, and tune it using evidence from the service you actually operate.