When people talk about sub-10-microsecond order routing, they often mean the headline number: tick arrives, order departs, wall clock says 6 microseconds elapsed. What gets less attention is everything that happens inside that window. The pipeline is not a single step. It is a sequence of discrete processing stages, each with a budget, and understanding where time goes is the first step toward tightening it.
This post walks through the stages we instrument in the EurekaLabs routing core, what each stage actually does, and where latency tends to accumulate in practice.
The Pipeline at a Glance
A routing decision involves five sequential stages once a market data update arrives at the feed handler:
- Feed handler decode and timestamp
- Order book state update
- Routing signal evaluation
- Venue selection and order construction
- Gateway serialization and kernel egress
The 6-microsecond figure we use in our lab benchmarks is the median across all five stages, measured from hardware timestamp on NIC receive to hardware timestamp on NIC transmit. It does not include network propagation time to the venue, which is a separate budget you own through your co-location choices.
Stage 1: Feed Handler Decode (0.4 to 0.8 us)
The first stage deserves careful treatment because it is where you set yourself up to succeed or fail in everything downstream. The raw market data packet arrives from the kernel (or directly from the NIC via kernel bypass) and needs to be decoded from the feed's binary format into the internal representation your routing logic understands.
For ITCH 5.0, the decode path is straightforward: fixed-width fields, no variable-length strings in the hot path, and every field you care about for routing (price, size, side, MPID) lands at a predictable byte offset. Our ITCH decoder runs in about 0.35 us at p50. CME MDP 3.0 uses SBE (Simple Binary Encoding) and adds a small schema-resolution cost, landing at roughly 0.5 us. The outlier in our environment is FAST-compressed feeds: decompression involves a state machine that does not vectorize cleanly, pushing decode time to 0.8 to 1.1 us for complex message types.
The practical implication: if you are handling multiple feed formats, your per-feed normalization cost is not uniform. This is worth measuring per message type, not just per protocol.
Stage 2: Order Book State Update (0.6 to 1.4 us)
Once decoded, the event needs to be applied to the in-memory order book. This is where most teams underestimate complexity. A clean add/modify/delete on a lightly loaded book is fast: a price-level lookup in a sorted container, an integer increment, done. But realistic conditions introduce variability.
Two sources of non-trivial cost: first, price level cache misses when a trade sweeps through multiple levels, forcing cache lines that were cold. Second, top-of-book recalculation when the best bid or offer changes. If your routing logic only cares about NBBO, you can skip rebuilding deeper levels, which is an optimization we apply selectively based on the venue's typical order flow profile.
The order book update stage is also where sequencing matters. If you are consuming redundant feeds (primary plus backup), you need a duplicate-suppression mechanism here. Badly designed dedup logic can add more latency than the redundant feed saves in failover scenarios. We cover this in more depth in a separate post on redundant feed architecture.
Stage 3: Routing Signal Evaluation (0.8 to 1.5 us)
This is the most variable stage and the one where architectural choices have the largest spread in outcome. The routing signal is a scalar score for each candidate venue that summarizes current spread, queue depth, recent fill rate, and any adverse selection signal from the feed. A simpler implementation might be a weighted sum of pre-computed venue stats. Our production routing model adds a few non-linear terms to handle venue misbehavior detection, which costs measurable additional time.
What matters most here is branch prediction and memory access pattern. The venue scoring table needs to be in L1 cache when the signal evaluation runs. We pin the routing thread to a dedicated CPU core and use prefetch hints to pull the venue stats structure into cache before it is needed. Without this, a cold venue stats fetch can add 0.4 us of memory latency by itself.
We are not suggesting every team needs the same complexity. If you trade at a single venue with minimal competition for queue position, your scoring function can be trivial and your signal evaluation budget drops to 0.2 to 0.3 us. The point is that routing logic complexity has a direct and measurable cost that you should understand before adding it.
Stage 4: Order Construction (0.3 to 0.6 us)
Once the target venue is selected and allocation is computed, the order needs to be constructed in the format the venue gateway expects. For FIX sessions this means serializing tag-value pairs into a pre-allocated buffer. For binary protocols it means writing fields at fixed offsets into a struct.
The cost here is dominated by how well you avoid dynamic memory allocation. Any heap allocation in the hot path is a latency spike risk because allocators have non-deterministic behavior under load. We pre-allocate a pool of order structs at startup and reuse them from a lock-free ring buffer. The allocation cost on the hot path is then a single atomic increment, which is negligible.
Order construction also handles FIX sequence number management and checksum computation for FIX sessions. Checksum computation over a 200-byte FIX message takes roughly 80 to 120 nanoseconds with SIMD acceleration. Without it, you are closer to 400 ns. This is one of the unglamorous optimizations that adds up.
Stage 5: Kernel Egress and Hardware Timestamp (0.5 to 1.2 us)
The final stage is getting the constructed order out of the process and onto the wire. With a standard kernel socket, the system call overhead plus TCP/IP stack processing runs 5 to 20 us by itself, which blows the entire routing latency budget. This is why kernel bypass is non-negotiable for sub-10-us routing.
With DPDK or RDMA, the egress path bypasses the kernel entirely. The outgoing packet is written directly to the NIC's transmit ring by user-space code. Hardware timestamps on the transmit side are what we use as the stop marker for our tick-to-trade measurement. The egress stage with kernel bypass runs 0.5 to 1.2 us depending on packet size and NIC queue depth at the moment of transmission.
Where Optimization Attention Is Worth Spending
Across the five stages, the ranking by variance in our test environment is: routing signal evaluation (most variable), order book state update, gateway egress, feed decode, order construction (least variable). High variance stages are where p99 latency diverges most from median, and p99 is usually what matters for competitive execution scenarios.
We want to be clear that these numbers reflect our testbed: specific NIC hardware, specific CPU generation, specific memory layout, specific feed formats. Translating them to your infrastructure requires running the same instrumentation on your stack. The methodology for that measurement is the subject of a separate post.
What the breakdown does tell you, regardless of your specific hardware, is where to look first when you are trying to understand a latency regression. Feed protocol change and order book path are the two most common culprits in our experience. A new feed format that adds 0.3 us to the decode stage will show up as a 5% increase in your median routing latency and a larger increase in your p99. Knowing which stage the regression is in tells you which team owns the fix.