Getting order routing below one millisecond is not a single optimization. It is the result of applying a set of architectural patterns consistently across the entire pipeline. Each pattern addresses a specific source of latency that would be above the budget without it. Applying four of the five and skipping one usually means you are not below the threshold, because the remaining source typically contributes enough to push you over.
This post covers the five patterns we consider foundational for sub-millisecond order routing. Each is well-documented in systems literature, but understanding how they interact in a complete routing pipeline requires seeing them together.
Pattern 1: Lock-Free Single-Producer Single-Consumer Queues
The classical mutex-based queue is not a viable option for passing events between the feed handler and the routing engine at low latency. A mutex contention under load can add tens of microseconds of scheduling latency, and even an uncontested mutex acquire costs around 50 to 100 ns due to the memory barrier it implies. For a pipeline targeting 500 us total budget, that is unacceptable.
The solution is a lock-free ring buffer sized to a power of two, shared between one producer thread (the feed handler) and one consumer thread (the routing engine). The producer writes to the head pointer; the consumer reads from the tail pointer. Neither ever blocks. Coordination is purely via atomic loads and stores with appropriate memory ordering constraints (release on write, acquire on read).
The critical implementation detail is cache line alignment. If the head and tail pointers share a cache line, updates to one invalidate the other in the other CPU's cache, creating false sharing that generates cache miss latency on every operation. Head and tail must be on separate 64-byte cache lines. This is one of those bugs that is invisible in single-threaded testing and catastrophic in production under load.
Pattern 2: Kernel-Bypass Networking
The Linux kernel TCP/IP stack is optimized for throughput and compatibility, not for minimum latency. For raw socket operations, the path from user-space write to NIC transmit involves a system call, kernel buffer management, TCP state machine processing, and NIC DMA. This typically runs 5 to 15 us on modern hardware, which consumes most or all of a sub-10-us routing budget just for the network egress step.
Kernel bypass moves the NIC interface into user-space entirely. DPDK is the most widely deployed framework for this on Linux: it provides a poll-mode driver that maps NIC transmit and receive rings directly into user-space memory. Your application writes outgoing packets directly to the transmit ring and polls the receive ring for incoming packets, with no system calls and no kernel involvement.
The cost of DPDK is operational complexity: you need a dedicated CPU core for the polling loop (a core that is 100% busy even when there are no packets, because polling), BIOS configuration for huge pages and NUMA-aware memory allocation, and vendor-specific driver setup. RDMA (specifically RoCE or InfiniBand) offers even lower latency for point-to-point connections but requires network infrastructure support and is typically available only in co-located environments.
Pattern 3: CPU Pinning and NUMA Isolation
Modern multi-socket servers have non-uniform memory access: memory on the remote NUMA node takes 40 to 100 ns longer to access than memory on the local node. For a pipeline where individual stages are budgeted in hundreds of nanoseconds, a NUMA remote access in the hot path is a material cost.
CPU pinning means assigning specific threads to specific CPU cores and preventing the OS scheduler from migrating them. The routing engine thread and the feed handler thread should both be pinned to cores on the same NUMA socket as the NIC that is receiving feed data and sending orders. The NIC's DMA writes go to the local socket's memory, and the threads that consume that memory pay local-access latency.
Beyond NUMA, the OS scheduler introduces jitter even without migration: a thread may be preempted to run a kernel task, a timer interrupt, or a higher-priority process. For the routing thread, even a 10 us preemption is visible in the p99 latency distribution. The standard mitigation is setting the thread's scheduling policy to SCHED_FIFO with a high priority, which reduces (but does not eliminate) involuntary preemption. Fully isolating the CPU with isolcpus kernel parameter eliminates it almost entirely at the cost of reducing the pool of cores available to the OS for general-purpose work.
Pattern 4: Pre-Computed Venue Scoring Tables
If your routing decision involves computing a score for each candidate venue at order arrival time, you have a budget problem. At 50 us total budget, a scoring function that does division, floating-point operations, or conditional branches on each of 15 venues can easily consume 5 to 10 us by itself.
The alternative is maintaining a pre-computed score table: a fixed-size array indexed by venue ID where the value at each index is the current score for that venue. When a routing decision is needed, the selection is just an argmax over a small integer array, which runs in tens of nanoseconds.
The scoring function itself runs continuously in a separate lower-priority update thread, consuming market data events and writing updated scores to the table. The routing thread only reads the table; the update thread only writes. With proper memory ordering on the writes (store-release) and reads (load-acquire), the routing thread always sees a consistent value that is at most one market data event stale. For most routing strategies, that staleness is negligible relative to the network propagation latency between your infrastructure and the exchange.
Pattern 5: DPDK-Accelerated FIX Serialization
FIX serialization into a pre-allocated transmit buffer is not inherently slow, but the combination of tag-value encoding, checksum computation, and length field back-filling can accumulate latency if not implemented carefully. The reference implementations that come with most FIX engines are not optimized for minimum latency, because general-purpose FIX engines target throughput and reliability, not microsecond-level serialization speed.
Key optimizations for FIX serialization in a latency-sensitive context: write fields to a pre-allocated buffer at fixed offsets (no dynamic resizing), compute the checksum with SIMD (SSE2 or AVX2 byte summation over the message buffer), and maintain a persistent FIX session state object that is pre-initialized so that session establishment fields (BeginString, SenderCompID, TargetCompID) are already written and do not need to be re-serialized on each order. The variable fields per order are OrderID, ClOrdID, price, quantity, and side. Serializing just those fields plus the checksum update runs under 200 ns with SIMD acceleration.
Why All Five Together
Each pattern eliminates a specific bottleneck. Skipping any one typically means one stage in your pipeline is an order of magnitude slower than the others, pushing the total above the sub-millisecond threshold. A pipeline with lock-free queues but standard kernel sockets is bound by the kernel networking cost. A pipeline with DPDK but mutex-protected routing state is bound by lock contention under load. A pipeline with everything right but cross-NUMA memory access has unpredictable p99 behavior.
The patterns work because they address independent bottlenecks. Combining them is multiplicative in effect, not additive. The order in which you implement them matters less than the completeness of the implementation.