Back to Pulse
EurekaLabs Pulse

Building Execution Infrastructure That Keeps Running When Components Fail

Haruto Yamane
Building Execution Infrastructure That Keeps Running When Components Fail

Resilience Is Not the Same as Redundancy

The words are often used interchangeably, but they point to different design goals. Redundancy means having backup copies of components. Resilience means the system continues to function correctly when components fail, which may or may not require redundancy depending on the component and the failure mode.

A trading infrastructure stack that maintains redundant feed connections but has a single-point-of-failure routing engine is redundant in one layer and fragile in another. Designing for resilience requires thinking about each layer independently: what failure modes exist here, what is the observable impact of each failure mode, and what is the correct system behavior in response?

Feed Layer: Graceful Degradation vs Hard Failover

A feed connection can fail in several ways: the physical connection drops and the OS detects it quickly, the feed produces sequence gaps indicating message loss without a connection drop, or the feed produces messages but they are stale (the exchange side is buffering and the data is no longer current). Each of these requires a different response.

For complete connection loss, hard failover to a secondary feed is appropriate. If you have a backup connection to the same venue, promote it to primary. If you do not, mark that venue as unavailable in the routing engine and remove it from active consideration until the connection recovers.

Sequence gaps are subtler. A small, short gap during a high-activity period may self-heal quickly if the retransmission channel catches up. A gap that widens over more than 2-3 seconds indicates a more serious data delivery problem. The right response to widening gaps is to treat the feed as degraded and reduce its weight in the order book reconstruction, rather than either ignoring the gap or immediately failing over. Failing over too aggressively for transient gaps causes its own problems: connection flapping and unnecessary state reconstructions.

Feed staleness without sequence gaps is the hardest failure mode to detect. If an exchange is buffering messages on their side, your feed will arrive with valid sequence numbers but timestamps that indicate the data is not current. The detection requires comparing exchange-side event timestamps (when available in the protocol) against your local receive time. A persistent offset larger than your configured threshold should trigger the same degraded-weight response as sequence gaps.

Order Book Layer: State Corruption and Recovery

Order book state corruption is a silent failure mode. The book appears intact but reflects an incorrect state because a message was dropped, processed out of order, or handled incorrectly during a prior failure event. Detecting this requires periodic reconciliation: either against snapshot updates from the exchange (when the protocol provides them) or against an independently-maintained second book.

The standard approach for ITCH-based feeds is to maintain a sequence-number counter and trigger a TCP retransmission request immediately when a gap is detected. The retransmission gap fill mechanism replays the missing messages in order, allowing the book to catch up without a full state reset. The risk is that the retransmission channel has its own latency and the retransmitted messages arrive slightly out of order with concurrent real-time messages. Your merge logic needs to handle this: apply retransmitted messages based on sequence number, not arrival time.

For failure modes that corrupt state beyond what gap fill can repair (corruption due to a bug, not a delivery gap), the only correct recovery is to discard the current state entirely, disconnect, reconnect, wait for a snapshot (or replay from the start of session), and rebuild. This is slow. Any design that avoids it entirely by building in correctness guarantees earlier in the stack is worth the investment.

Routing Engine: Stateful Failover

The routing engine is the hardest component to make resilient because its state is significant: real-time venue scores, open order tracking, session-level fill statistics, and current market context. If the primary routing engine process crashes, promoting a standby means the standby needs to have approximately current state, not just a cold-started version that has to rebuild from scratch.

Two approaches work in practice. The first is state replication: the primary routing engine publishes state updates to a shared-memory segment or low-latency IPC channel, and the standby consumes them continuously. On failover, the standby is already warm and can pick up routing within milliseconds. The cost is the CPU and memory overhead of state replication on the hot path.

The second approach is to design the routing state so that it recovers quickly from cold start. Venue scores that update on a rolling window of recent data reach an accurate steady state within a few seconds of receiving live feed data. If you can tolerate routing decisions at default-score quality for the first 5-10 seconds after a failover, warm standby state replication may not be necessary for your use case.

Order Gateway: Session Continuity

Order gateway failures during active trading create the most immediately visible execution problems: open orders that are neither cancelled nor filled, FIX session sequence number state that needs reconciliation, and potential double-submission if the failover process is not careful about what was already sent.

The cardinal rule for order gateway resilience is that order submission must be idempotent at the infrastructure level. The gateway must maintain a record of every submitted order message and its submission status before considering it sent. On failover, the new gateway instance reconciles its open order state against the exchange's order status before submitting anything new. This prevents double submission, which creates both execution risk and compliance issues.

FIX session failover requires particular care because FIX session sequence numbers are strictly ordered and exchange-side desynchronization triggers a session-level resync that halts trading for several seconds. Maintaining FIX session state in a shared store accessible to both primary and standby gateway processes is the standard solution.

Testing Resilience: The Part That Gets Skipped

The most common gap in resilience engineering is not in the design but in the testing. Failover mechanisms that look correct in code and in architecture diagrams fail in production because they have never been exercised under realistic conditions.

The minimum viable resilience test suite: periodic forced failover of each redundant component during a non-trading period, to verify that the failover path actually works. Feed gap injection, to verify that gap detection and recovery work correctly and do not corrupt state. Order gateway session reset testing, to verify that sequence number reconciliation works after a restart.

Testing these in production is uncomfortable. Not testing them means the first time your failover runs is during a real incident. For trading infrastructure, that tradeoff is usually easy to resolve: the cost of a planned failover drill is far lower than the cost of an unplanned production failure.

EurekaLabs

See the Infrastructure Behind These Numbers

Request access to the EurekaLabs platform and run your own latency benchmark against your current infrastructure.