Every vendor in the low-latency trading infrastructure space publishes a latency number. Ours is on our website. The problem is that latency numbers mean nothing without a methodology specification, and methodology specifications are rarely published in enough detail to be comparable across vendors or to reproduce in an independent test environment.
This post is our methodology documentation. We describe exactly how we measure tick-to-trade latency, what hardware we use to timestamp, what test conditions produce the numbers we publish, and what the numbers do not tell you. The goal is to give you enough information to reproduce our measurement on your own infrastructure if you want to verify or compare.
What Tick-to-Trade Latency Measures
Tick-to-trade latency measures the interval between a market data event arriving at your system and the corresponding order reaching the exchange's matching engine network boundary. The "tick" end is when a specific market data packet arrives at your infrastructure. The "trade" end is when the order you generate in response to that packet is transmitted onto the network toward the exchange.
Note what is not included: network propagation time from your infrastructure to the exchange, exchange processing time, and execution latency at the exchange. Those are separate from your system latency and outside your control once you have chosen your co-location strategy. Tick-to-trade measures only the latency your system introduces.
This distinction matters because marketing materials sometimes blend these components. A "sub-millisecond round-trip" claim may include one-way network propagation at 50 to 150 us plus exchange acknowledgment at 50 to 200 us, which means the actual system processing time could be 300 to 700 us and still satisfy the headline claim. The headline is true but the meaning is different from what "sub-millisecond order routing" implies.
Hardware Timestamping: Why Software Clocks Are Not Enough
Software timestamps (from the OS clock, accessed via clock_gettime) introduce measurement error from two sources: interrupt jitter and OS scheduling preemption. When your process calls clock_gettime, the returned value reflects when the kernel handled the call, not necessarily when the underlying event occurred. For events that arrive via interrupt (which includes most network events on a standard kernel), the timestamp can be off by the interrupt latency of the system, which varies from nanoseconds to tens of microseconds depending on CPU load and kernel configuration.
Hardware timestamps are assigned by the NIC at packet arrival time, based on a hardware clock that runs independently of the kernel. The NIC hardware clock is synchronized to PTP (Precision Time Protocol) or GPS-locked timing to maintain accuracy relative to an external reference. For transmit, the hardware timestamp records when the packet actually left the NIC. Neither timestamp involves the kernel or user-space code, so they are free of software jitter.
In our testbed, we use a Mellanox ConnectX-6 NIC with hardware timestamp support on both receive and transmit paths. The NIC hardware clock is disciplined with PTP against a GPS-referenced grandmaster clock in the data center. Our measurement code reads the hardware timestamps from the NIC after each event and stores them in a ring buffer that the analytics process reads offline.
Test Environment Specification
Processor: Dual Intel Xeon Gold 6338 (Ice Lake, 3.2 GHz boost), 32 cores per socket. Routing thread pinned to physical core 4, NUMA socket 0 (same socket as NIC). BIOS: hyperthreading disabled, C-states disabled, frequency scaling disabled (performance governor). Turbo boost enabled; frequencies locked to maximum.
Memory: 256 GB DDR4-3200 ECC, interleaved across DIMM slots on NUMA socket 0. Huge pages: 2 MB pages pre-allocated for DPDK memory pool at system startup.
Network: Mellanox ConnectX-6 25 GbE NIC, DPDK PMD, connected via physical cross-connect to NYSE Pillar market data feed endpoint and NYSE order entry endpoint in NY4. Kernel bypass for both receive (market data) and transmit (orders) paths.
Operating system: Ubuntu 22.04 LTS, kernel 5.15 with PREEMPT_RT patch. CPU isolation via isolcpus for cores 4 and 5. Routing engine thread on core 4, feed handler on core 5. IRQ affinity: all NIC interrupts bound to core 3.
Software: EurekaLabs routing core with NASDAQ ITCH 5.0 feed handler, FIX 4.4 order serializer, pre-computed venue scoring table. No other processes on isolated cores during test runs.
Test Procedure
We replay a captured NASDAQ ITCH stream recorded during a typical session (approximately 1.8 million messages per minute at mid-session). The replay feeds packets into the DPDK receive ring at the original inter-packet timing, using hardware timestamps from the capture to maintain realistic inter-arrival patterns. For each message that triggers a routing decision (we use a simple signal: any top-of-book update to the five instruments in the test set), we record the receive hardware timestamp and the transmit hardware timestamp on the resulting order.
The test runs for one hour. We collect approximately 14,000 routing decisions per run. We repeat each test configuration three times and use the second and third run for the reported statistics (first run is treated as warmup for caches and branch predictors).
Why Median Numbers Are Misleading
The median tick-to-trade latency in our testbed is around 4 to 5 us. This is the number that looks good on a data sheet. It is also the least useful number in the distribution for understanding operational behavior.
Trading strategies that depend on sub-millisecond execution care about tail latency because execution quality degrades nonlinearly with latency. A routing decision that takes 50 us instead of 5 us may arrive after the market has moved against you, turning a neutral or positive expected value execution into a negative one. The fraction of orders in the tail determines how often this happens.
The statistics we publish alongside median: p95, p99, p99.9, and the maximum observed in each test run. The distribution from our testbed: median 4.8 us, p95 7.2 us, p99 9.1 us, p99.9 22 us, max observed 68 us across three one-hour runs. The max-observed number is the most informative for worst-case planning.
The p99.9 and max numbers are where differences in OS configuration and hardware show up most. The same software running on a non-isolated CPU with a standard kernel scheduler shows a p99 of 40 to 80 us and occasional spikes above 500 us due to kernel timer interrupts and scheduler preemption. The OS configuration work is what separates a software-latency claim from a system-latency claim.
What the Numbers Do Not Tell You
Our testbed numbers are from a controlled replay environment with a clean DPDK receive path, no competing processes, and a hardware-disciplined clock. Production environments differ in ways that matter: other processes on the same physical host (even on non-isolated cores) generate LLC pressure and NUMA traffic. Real feed arrival patterns under stress (quote-stuffing events, market-open rush) differ from the mid-session replay. And network round-trip to the exchange adds 50 to 150 us that our tick-to-trade measurement does not include.
We publish these numbers as what they are: the latency budget that our software and hardware contributes to the pipeline, measured under controlled conditions. They are a starting point for understanding your system latency, not a promise about end-to-end execution quality in production.