Architecture analysis · Research date: October 3, 2026 · Includes a runnable Java 21/JDBC example and 13 passing tests.

A support agent decides that a customer should receive a €25 credit. The receiving service commits the credit. Then the worker disappears before it reports success.

The workflow sees an unfinished step. The customer already has the money.

What should happen when the workflow retries?

This is where an impressive agent demo becomes a distributed systems problem. Saving a conversation or restarting a worker does not settle whether a business action already happened. Durable orchestration preserves progress. Safe effects require their own contract.

Why this is a timely architecture topic

Homann Software's focus on Java, Spring, domain-driven design and iterative engineering makes production AI reliability a natural area to examine. The current primary-source landscape offers several relevant directions: more selective RAG pipelines, typed model decisions, modern Java concurrency, and durable execution for agents. Homann's engineering profile connects these topics to practical application architecture.

The freshest Spring signal is the October 2 article on Modular RAG and TypeSafe Jev, which combines retrieval with a relevance decision stage. The September 25 Spring AI 2.1 milestone also shows active development. A milestone remains a pre-release; its announcement alone is no basis for a production upgrade.

For this article, we selected durable agent execution because it addresses a different gap in our recent coverage: what happens after a typed, authorized tool proposal reaches a system that changes business state? This is an editorial choice based on relevance, recency and practical testability, rather than a claim that one topic dominates the entire industry.

On September 21, Temporal announced a prerelease integration using Amazon Bedrock AgentCore Runtime as a compute provider for Serverless Workers. Its architecture separates durable execution state from worker lifetime and still requires safe external effects. The published reference sample is Python; this article does not imply a tested Java implementation of that integration. Temporal's announcement is a useful current signal, not independent evidence of market adoption.

Separate proposals, process state and committed effects

An AI component can propose a credit. A domain policy decides whether that credit is allowed. A workflow coordinates the steps and records their outcomes. The receiving service owns the actual state change.

Those responsibilities need separate evidence:

Boundary Evidence to retain Question it answers
Model proposal Typed output and model/prompt revision What did the model suggest?
Domain authorization Actor, validated command, approval revision Was this particular action allowed?
Workflow history Scheduled steps, recorded results, deadlines Where should execution resume?
Receiver receipt Operation identity, exact payload, committed result What actually changed?

Architecture diagram separating AI proposal, authorized workflow command and receiver transaction.

A model's proposal is not a receipt. The diagram describes an intended application architecture; the executable example below covers its receiver transaction.

In Temporal, orchestration must remain deterministic during replay. Model calls, database access and other external interactions belong in Activities. Replay can use recorded Activity results instead of asking the model to make the same completed decision again. New or retried calls can still produce different outputs, so bind accepted decisions to an explicit command revision. Temporal's Workflow Definition documentation explains the replay constraints.

This separation matters even without AI. An agent simply makes the number and variety of external interactions more visible.

The dangerous interval is after commit, before confirmation

There are two distinct failure cases.

If the receiver's transaction never committed, retrying can create the first effect. If it committed but the response was lost, retrying must recover the existing receipt.

Sequence showing a committed credit, a lost acknowledgement and a retry with the same operation key returning the existing receipt.

Failure of the response is not evidence that the action failed.

Temporal explicitly documents a worker completing an Activity and crashing before notifying the service. That Activity can be retried even though its external effect happened. The idempotency guarantee therefore belongs at the service receiving the request. Activity Definition: idempotency describes this failure window.

The practical contract is:

  1. Create an operation identity when the application accepts the business command. Persist it outside the model, and reuse it on every delivery attempt.
  2. Scope the identity to the authenticated tenant and operation type. A revision such as refund-42-v1 identifies one accepted action, not every refund for an order.
  3. Bind that identity to the full effect-bearing payload. Changing the destination, amount or currency must produce a conflict.
  4. Atomically commit the receiver's local effect and its receipt.
  5. Return the saved receipt for an identical repeated request.

Generating a fresh random key on each retry defeats the contract. So does trusting the model to choose the key. Two distinct keys for one business action can still create two effects: ingress deduplication and any business rule limiting total credits remain separate responsibilities.

A runnable Java receiver

The example models a local account-credit ledger. It uses Java 21, JDBC and file-backed H2 for the process-recovery tests. It is a small receiver implementation, not a replacement workflow engine or payment service.

The caller supplies an already authorized Command: tenant, operation key, account, positive minor-unit amount and currency. In a real endpoint, derive tenant identity from authentication and enforce current domain policy before invoking the ledger. The example does not implement authentication or approval storage.

Two tables hold account credits and receipts. A composite primary key on (tenant, op_key) arbitrates concurrent duplicate requests. The following method is extracted directly from the compiled source:

    public String credit(Command cmd, Fault beforeCommit, Fault afterCommit) throws SQLException {
        try (var c = connect()) {
            c.setAutoCommit(false);
            String receipt = UUID.randomUUID().toString();
            try {
                try (var s = c.prepareStatement("INSERT INTO receipts VALUES(?,?,?,?,?,?)")) {
                    s.setString(1, cmd.tenant()); s.setString(2, cmd.key());
                    s.setString(3, cmd.account()); s.setLong(4, cmd.cents());
                    s.setString(5, cmd.currency()); s.setString(6, receipt);
                    s.executeUpdate();
                }
            } catch (SQLException e) {
                c.rollback();
                if (!"23505".equals(e.getSQLState())) throw e;
                return existingReceipt(c, cmd); // compare payload before returning
            }
            try {
                try (var s = c.prepareStatement("UPDATE accounts SET credit=credit+?"
                        + " WHERE tenant=? AND account=? AND currency=?")) {
                    s.setLong(1, cmd.cents()); s.setString(2, cmd.tenant());
                    s.setString(3, cmd.account()); s.setString(4, cmd.currency());
                    if (s.executeUpdate() != 1) throw new SQLException("Account/currency mismatch");
                }
                beforeCommit.run();
                c.commit(); // receipt and local effect become durable together
            } catch (SQLException | RuntimeException e) {
                c.rollback();
                throw e;
            }
            afterCommit.run(); // simulate an ambiguous result after a successful commit
            return receipt;
        }
    }

The insert comes first. A concurrent duplicate cannot independently pass an application-level “does this exist?” check and then apply another effect. The database enforces uniqueness. The credit update and receipt live in the same transaction; a rejected destination or a failure before commit rolls back both.

For a duplicate, the saved payload must match before the receipt is returned:

    private String existingReceipt(Connection c, Command cmd) throws SQLException {
        try (var s = c.prepareStatement("SELECT account,cents,currency,receipt FROM receipts"
                + " WHERE tenant=? AND op_key=?")) {
            s.setString(1, cmd.tenant()); s.setString(2, cmd.key());
            try (var r = s.executeQuery()) {
                if (!r.next()) throw new SQLException("Conflicting receipt unavailable");
                if (!cmd.account().equals(r.getString(1)) || cmd.cents() != r.getLong(2)
                        || !cmd.currency().equals(r.getString(3))) {
                    throw new IllegalArgumentException("Idempotency key reused with different payload");
                }
                return r.getString(4);
            }
        }
    }

Notice that the SQL exception handler recognizes the specific duplicate-key SQLSTATE used in this H2 example. Other database engines and frameworks can require different error handling. Review the semantics of the production database rather than copying a dialect-specific branch unchanged. H2 documents its transaction and concurrency behavior.

The fault hooks make the otherwise tiny failure window reproducible. beforeCommit can interrupt the transaction; afterCommit can lose the response after a successful commit. Ordinary callers pass RefundLedger.NONE for both hooks.

Test the uncertainty, not just the happy path

The retained run used Temurin OpenJDK 21.0.11+10 and H2 2.2.224, pinned for reproduction. No preview features, model subscription or provider credentials are needed.

The suite checks repeated requests; changed amount, destination and currency; tenant-scoped identities; invalid input; a missing account; rollback before commit; a lost response after commit; concurrent duplicate and distinct commands; and process termination on both sides of commit.

Failure matrix showing expected initial state and retry outcome for rollback, lost response and process termination around commit.

The matrix summarizes the controlled tests. It does not show a production availability measurement.

For the two file-backed recovery tests, a separate JVM calls Runtime.halt(73). The parent verifies the exit code, opens the database again, inspects the initial balance and receipt count, then retries the same command. H2 uses WRITE_DELAY=0 in these tests. A halt after commit leaves one credit; a halt before commit leaves none. Retrying converges to one receipt and €25 in both cases.

There is also a concurrent test with twelve attempts sharing one key. All callers receive the same receipt and the balance increases once. Twelve distinct keys produce twelve credits without losing an update.

On Windows, from the extracted examples directory with the pinned H2 JAR present:

javac --release 21 -d out RefundLedger.java CrashWorker.java LedgerTest.java
java -cp "out;h2-2.2.224.jar" LedgerTest

The saved output ends with:

PASS process halt before commit recovers with zero initial effect
PASS process halt after commit recovers with one initial effect
13 tests passed

Download the source and reproduction instructions. The README includes the pinned dependency download and Linux/macOS classpath variant. The JAR is not bundled in the download.

These are receiver-contract tests. They do not establish Temporal replay compatibility, a Spring transaction configuration, multi-host recovery, power-loss durability, external API behavior or payment correctness. The credit ledger also omits refund eligibility, maximum cumulative refunds and accounting rules. Those need their own domain model and tests.

A database transaction cannot wrap a remote provider

The example's effect is local. If an Activity invokes an external payment API, writing a local receipt cannot atomically commit that provider's action.

Send the accepted operation identity to the provider's supported idempotency mechanism. Understand its retention window and payload-matching rules. For example, Stripe documents returning the saved status and body for repeated keys, comparing request parameters, and permitting a new request after a key is pruned. A retry policy that outlives the provider's retained identity needs another recovery strategy. Stripe's idempotent-request contract illustrates why adapters must follow provider-specific semantics.

If a provider offers neither idempotency nor authoritative lookup by operation identity, an ambiguous response may require reconciliation or human investigation. Blind retry can duplicate the action; blind compensation can create a second error. Reconciliation is a valid workflow state, with an owner and evidence, not just a log message.

An outbox helps when a local transaction must schedule a later message. It can commit intent alongside domain state, but downstream delivery can still repeat. Consumers need their own deduplication or idempotency contract.

Bring the contract into an iterative delivery process

Start with one valuable use case, such as an approved account credit. Make success, rejection, changed payload and unknown outcome visible in the use-case model and UX. Let users distinguish “pending confirmation” from “failed”; a timed-out UI should not invite a new operation identity for the same action.

Use an ADR to record who owns the operation identity, where the receipt is authoritative, how long identity records are retained, and how workflow or command revisions are handled. Define deadlines, retry budgets and reconciliation ownership as quality requirements.

Then add failure injection to the delivery loop. Test old workflow histories before changing orchestration code; test adapters against the provider's actual contract; and test receipts and local state after ambiguous outcomes. Connect the acceptance scenarios to code, telemetry and operational procedures. Refine those artifacts together as feedback changes the design.

Durable execution is valuable when a process must wait, resume and coordinate work over time. Its value becomes concrete when every repeated tool call has a defensible answer to one question: what committed effect does this operation identity already represent?