Research date: October 2, 2026. Architecture guide with a tested Java 21 indexing core.

Your explorer shows a transfer. A user refreshes the page, and the transfer disappears. The transaction hash has not changed. The chain's canonical history has.

That is where a blockchain explorer becomes an interesting software architecture problem. A fast search box is useful. A trustworthy explanation of what happened, which block contained it, and whether that block still belongs to the chain is essential.

The current Ethereum ecosystem makes this a timely topic. Here is how to design the read model, test its failure paths, and decide how much explorer infrastructure you should build yourself.

Why this topic matters now

Three recent developments are worth an engineering team's attention:

  • Ethereum is preparing its next protocol upgrade. The Ethereum Foundation's announcement, dated September 28, schedules Glamsterdam on Sepolia for October 6, 2026, at 13:53:36 UTC. Its headline changes include enshrined proposer-builder separation and block-level access lists. Hoodi and mainnet dates remain undecided in that announcement. This is an upcoming testnet event, not a completed mainnet rollout. Ethereum Foundation announcement.
  • Explorer deployment is becoming easier. Blockscout introduced Autoscout Wizard on September 23, describing initial deployment using an instance name and RPC URL. That lowers the setup barrier and makes the build-versus-adopt decision more concrete. Blockscout's launch post.
  • Explorer privacy is receiving product attention. Blockscout's September 24 post describes its Tor-native Ethereum explorer, which it says has been running since February, with tracking and account features removed. The recent announcement should not be confused with its original deployment date. Blockscout's Onion Mode post.

Our engineering conclusion is that easy deployment leaves more room to focus on data correctness, provenance, and useful domain views. This is an editorial choice for Homann Software's Java, architecture, and blockchain focus, rather than a measured ranking of market popularity. The implementation below teaches durable indexing principles; it does not implement or validate Glamsterdam-specific RPC behavior.

An explorer is a read model of a changing chain

An Ethereum node provides protocol data. An explorer organizes that data for people and applications. Treat the explorer as a rebuildable projection with an explicit source of authority.

Start with a small use case: inspect a block and the events included in it. Specify the edge cases before expanding the UI: duplicate delivery, reconnects, conflicting histories, missing parents, stale notifications, and a provider reporting an inconsistent finalized checkpoint.

Architecture sketch showing RPC ingestion, branch reconciliation, canonical storage, and a query API

Figure 1. Proposed production architecture. The downloadable Java example implements the reconciliation and projection core; the RPC adapter, database, API, and UI are design extensions.

Keep three concepts separate in your model:

Concept Identity Purpose
Observed block Network identity + block hash Preserve the block and where it was observed, even if orphaned
Canonical block Network identity + height → block hash Describe the current selected history
Event occurrence Network identity + block hash + log index Identify an inclusion without mixing competing blocks

Use the chain ID plus a configured genesis hash to distinguish networks in production. Private environments can reuse chain IDs. The example uses a single configured chain ID and trusted genesis fixture, which is sufficient for its controlled test cases.

A transaction hash identifies a transaction, but an inclusion needs its block identity too. A height alone cannot distinguish competing blocks. Preserve raw observations and derive the current canonical view from them.

Ethereum's JSON-RPC API exposes block hashes, parent hashes, receipts and logs, along with latest, safe, and finalized block tags. Fetch receipts and logs for the selected block and check that their block identities agree before constructing a complete batch. Ethereum JSON-RPC documentation.

Reconnects require reconciliation

Geth documents that subscriptions deliver current events, are tied to a connection, and do not provide historical replay. Its newHeads stream can report multiple headers at the same height during reorganizations. Log subscriptions can resend old occurrences with removed: true. Geth real-time events documentation.

Use a notification to wake the indexer. The indexer should then read its configured provider's current head, fetch ancestors by hash until it reaches a retained canonical block, and construct a contiguous replacement suffix. Bound that walk and fail visibly when the ancestor is outside your retained window.

Recheck the selected head before committing. If the provider changes its view during collection, retry the collection rather than combining two histories. On reconnect, repeat reconciliation from persisted state. Avoid advancing a cursor solely because a WebSocket notification arrived.

Those steps are our proposed adapter contract, not features exercised by the dependency-free example. They should receive separate integration tests with a controlled RPC service and then a supported node.

Replace the suffix, including its projections

Consider this sequence:

  1. The explorer has indexed G → A1 → A2.
  2. Its trusted provider now selects G → A1 → B2 → B3.
  3. A1 is the common ancestor. A2 becomes an orphaned observation.
  4. The canonical mapping and its events change together.

Reorganization sketch with A1 as common ancestor, A2 orphaned, and B2 and B3 replacing the suffix

Figure 2. Keep the old block as an observation, remove its events from the current projection, and install the replacement suffix.

Here is the central operation from the runnable Java example:

public synchronized boolean reconcile(List<Block> candidate) {
    var branch = List.copyOf(candidate);
    if (branch.isEmpty()) throw new IllegalArgumentException("empty branch");
    var first = branch.getFirst();
    var parent = canonical.get(first.number() - 1);
    if (parent == null || !parent.hash().equals(first.parentHash()))
        throw new IllegalArgumentException("unknown canonical ancestor");

    var hashes = new HashSet<String>();
    for (var block : branch) {
        if (block.chainId() != chainId) throw new IllegalArgumentException("wrong chain");
        if (block.number() != parent.number() + 1 ||
            !block.parentHash().equals(parent.hash()))
            throw new IllegalArgumentException("disconnected branch");
        if (!hashes.add(block.hash())) throw new IllegalArgumentException("reused hash");
        var known = observed.get(block.hash());
        if (known != null && !known.equals(block))
            throw new IllegalArgumentException("conflicting block contents");
        parent = block;
    }
    var tip = branch.getLast();
    var currentAtTip = canonical.get(tip.number());
    if (tip.equals(currentAtTip)) return false;
    if (first.number() <= finalized.number())
        throw new IllegalStateException("branch crosses finalized checkpoint");

    // Validate and build a replacement state before touching live state.
    var next = new TreeMap<>(canonical);
    next.tailMap(first.number(), true).clear();
    var nextObserved = new LinkedHashMap<>(observed);
    for (var block : branch) {
        next.put(block.number(), block);
        nextObserved.put(block.hash(), block);
    }
    var nextEvents = project(next.values());
    canonical = next;
    observed = nextObserved;
    events = nextEvents;
    return true;
}

The method accepts a branch immediately after a retained canonical ancestor. The caller must already have selected that branch through its configured provider. It does not choose the longest branch or reproduce Ethereum fork choice.

Validation happens before state replacement. A malformed batch leaves the state unchanged. Previously observed hashes cannot acquire different contents, and replaying an already canonical tip does not rewind a newer head. A shorter competing branch is accepted when the caller explicitly reconciles it as authoritative.

For this small in-memory example, synchronized access and copy-before-swap provide atomic visibility within one process. Rebuilding the event map also makes orphan removal easy to inspect. This is intentionally unsuitable for mainnet scale: retained observations grow without bound, and each update rebuilds the canonical projection.

In a database-backed implementation, put branch replacement, affected projection changes, and checkpoint updates in one transaction. Use a per-network writer or a fencing mechanism so two workers cannot commit incompatible branches. Store outbound change notifications in a transactional outbox; publish them after commit. External consumers need compensating updates when a previously reported inclusion is orphaned. These are production design requirements, not tested database guarantees of this sample.

Finality belongs in the data contract

A successful transaction receipt describes execution. Canonical inclusion describes the currently selected history. Finality provides a stronger consensus commitment. Ethereum's Gasper documentation explains finality and its relationship to fork choice; ordinary depth counting is a different policy. Ethereum Gasper documentation.

The example accepts a finalized checkpoint only when its height and hash match the local canonical chain. It rejects checkpoint regression and a replacement that crosses that checkpoint. If this invariant fails in production, stop ingestion for the affected network and investigate the provider and persisted state; do not silently rewrite the checkpoint.

Expose distinct UI states for a pending transaction, canonical inclusion, provider-reported safe inclusion, provider-reported finality, and an orphaned inclusion. Attach the indexed head hash and observation time to responses so users can see the view behind a result.

On an L2, also model the rollup-specific settlement stages. An L2 execution receipt alone does not establish L1 settlement. Define those meanings per supported network instead of copying Ethereum L1 badges into every explorer.

Run the Java example

The downloadable source package contains ReorgIndex.java, ExplorerDemo.java, ReorgIndexTest.java, and the optional mutation runner. It requires JDK 21 or newer and no Java dependencies. This is a runnable indexing core and CLI demonstration, not a complete browser-based explorer.

From the extracted examples directory:

javac --release 21 -d out ReorgIndex.java ExplorerDemo.java ReorgIndexTest.java
java -cp out ReorgIndexTest
java -cp out ExplorerDemo
python verify_mutations.py

Python 3 is needed only for the final command. The mutation runner writes disposable variants below mutation-work; it leaves the original Java sources untouched.

The demo applies the reorganization above a finalized A1 checkpoint:

var index = new ReorgIndex(block(0, 1, 0, "genesis"));
index.reconcile(List.of(block(1, 2, 1, "A1"), block(2, 3, 2, "A2")));
index.finalizeAt(1, hash(2));
index.reconcile(List.of(block(2, 4, 2, "B2"), block(3, 5, 4, "B3")));
var view = index.snapshot();

Its executed output is:

Canonical: [genesis, A1, B2, B3]
Observed blocks (including orphan): 5
Canonical events: 4
Finalized height: 1

The hexadecimal hashes are synthetic test identifiers, and the event strings are fixture values. The model validates hash shape and normalization, not Ethereum header hashing, receipt authenticity, signatures, or consensus proofs. A production decoder must validate the full RPC payload before it enters this core.

What the tests establish

Executed on Eclipse Temurin OpenJDK 21.0.11+10, the suite passed 19 named tests, including 108 combinations of original branch length, common ancestor, and replacement depth. These combinations are exhaustive within the stated test ranges, not across all possible Ethereum histories.

Four separately compiled, deliberately broken variants were all rejected by the unchanged tests:

Infographic mapping seeded indexing regressions to the tests that reject them

Figure 3. Local test evidence: all four seeded regressions compiled and failed the unchanged test suite. This is not a general mutation score or a performance measurement.

Seeded regression Required behavior
Remove the finality guard Reject a conflicting replacement below the checkpoint
Keep the old suffix Eliminate blocks beyond the replacement tip
Accumulate old events Exclude orphan events from the current projection
Accept conflicting contents under one hash Reject inconsistent observed block data

The tests also cover replay, stale canonical notifications, branch restoration, chain mismatch, unknown ancestors, malformed batches, duplicate event indices, and immutable snapshots. Raw results and the mutation report accompany the sources.

They establish behavior of this in-memory model under controlled inputs. They do not establish node interoperability, persistence after a crash, RPC recovery, multi-worker coordination, receipt completeness, malicious-provider resistance, Spring integration, throughput, or production availability.

Build the domain view; evaluate existing infrastructure

For a general-purpose EVM explorer, evaluate Blockscout before committing to a custom implementation. Its documentation describes node and tracing requirements; capabilities and historical data availability depend on the client configuration. Validate those against the exact client and provider you will operate. Blockscout client requirements.

A custom component is easier to justify when you need business semantics: a permissioned-chain audit view, domain-specific event search, a settlement workflow, or a narrowly scoped operational dashboard. An established explorer and a Java domain projection can coexist.

Turn that decision into a short ADR covering the supported networks, trust model, retention, query scope, and recovery requirements. Keep implementation iterative: build one useful inspection flow, test its edge cases, then add the next view using user feedback.

Before calling the service production-ready, test reconnect and backfill against a controlled RPC endpoint, crash recovery against the chosen database, and competing writers against the actual coordination mechanism. Monitor ingestion lag, reorganization depth, provider disagreement, and age of the last finalized checkpoint. Load tests must reflect the supported chains and query patterns.

A trustworthy explorer makes its view of the chain explainable. Start with stable identities, replace canonical suffixes coherently, preserve observations, and give users an honest description of finality. A polished interface is much more useful when those foundations hold.