Back to Pulse
EurekaLabs Pulse

Co-location vs Cloud: What the Latency Numbers Actually Show

Dmitri Sotnikov
Co-location vs Cloud: What the Latency Numbers Actually Show

The conversation about co-location versus cloud for trading infrastructure often generates more heat than light, because the two sides are not measuring the same thing. Co-location advocates cite round-trip times. Cloud advocates cite operational simplicity and global reach. Both are right about what they are measuring. The question is whether what they are measuring is what you actually need.

We ran the same routing workload from two environments: a 1U server in a NY4 cage with a cross-connect to NYSE Pillar, and an AWS us-east-1 instance (c6i.4xlarge) using a standard Ethernet connection to the same exchange's cloud-accessible order entry gateway. The results are worth sharing in detail because the gap is larger in some dimensions and smaller in others than most people expect.

Test Setup and What We Measured

The workload was a simplified version of our routing pipeline: receive a synthetic market data update, evaluate a venue scoring function against a static snapshot of order book state, construct and serialize a FIX 4.4 new order single, and write it to a socket. We measured three things: median tick-to-order-write latency, p99 latency, and variance across a one-hour session.

For the NY4 environment: dedicated server, Intel Xeon Ice Lake CPUs, Mellanox ConnectX-6 NIC with kernel bypass via DPDK, direct cross-connect to NYSE. We did not use co-located exchange matching engines for the test because we wanted to measure our own infrastructure, not exchange round-trip. The socket write in the co-location test goes to a FIX session handler that terminates at the cross-connect endpoint.

For the AWS environment: c6i.4xlarge, enhanced networking (ENA) enabled, standard kernel TCP/IP stack (no kernel bypass available in this environment), connection to the exchange's AWS-hosted order entry point via the exchange's connectivity offering from us-east-1. Same FIX session format, same order construction logic.

The Latency Numbers

Co-location (NY4, kernel bypass): p50 = 4.2 us, p99 = 9.1 us, p99.9 = 22 us.

AWS us-east-1 (standard kernel networking): p50 = 180 us, p99 = 420 us, p99.9 = 1.8 ms.

The gap at p50 is roughly 43x. The gap at p99 is roughly 46x. The gap at p99.9 is roughly 82x. Variance is dramatically higher in the cloud environment, which is expected: the kernel TCP stack, the hypervisor scheduling jitter, and the shared network fabric all contribute non-deterministic overhead that does not exist in a dedicated hardware environment with kernel bypass.

There is an important caveat: we did not enable DPDK or RDMA in the AWS environment because those are not available on standard EC2 instance types. AWS does offer placement groups and enhanced networking that narrow the gap somewhat, and for non-HFT workloads, the p50 cloud number may be acceptable. But for any strategy that competes on queue position or that depends on sub-millisecond reaction to market events, the cloud numbers are disqualifying.

When Cloud Is the Right Choice

We want to be direct about where cloud works well for trading infrastructure, because the co-location case is not always the right answer.

If your latency requirement is in the milliseconds range, not the microseconds range, cloud infrastructure is competitive. A lot of options market making, statistical arbitrage running on daily or minute bars, and systematic strategies with execution windows measured in seconds can tolerate 200 to 500 us of end-to-end routing latency without material impact on strategy performance. For these workloads, the operational benefits of cloud (elasticity, managed infrastructure, simplified deployment) are real and valuable.

Cloud also makes sense for non-production components of a trading stack. Backtesting infrastructure, analytics pipelines, parameter optimization runs, and risk aggregation are all well-suited to elastic cloud compute. The mistake is using cloud for components that touch the live execution path when your strategy cares about execution latency.

A hybrid architecture is often the right answer: co-location for the feed handler, routing core, and order gateway; cloud for everything else. The co-located components handle latency-sensitive execution; the cloud components handle scale and flexibility for everything that is not on the critical path.

Co-location Operational Costs Are Real

The co-location argument has a cost side that is worth being honest about. A NY4 cage requires a physical presence: server hardware, network hardware, cross-connects (which are billed per connection per month at rates that add up quickly when you are connecting to multiple venues), a colocation contract with the data center, and operational staff time to manage hardware incidents. Remote hands are available but expensive, and a hardware failure at 2am during a live trading session is not a comfortable scenario.

For an early-stage trading team, the operational overhead of co-location is not trivial. Cloud or a managed infrastructure provider can absorb that operational cost in exchange for higher per-unit pricing and worse tail latency. Whether that tradeoff makes sense depends on the latency sensitivity of your specific strategy and the stage of your build-out.

The Cross-Connect Cost That People Miss

One cost that often surprises teams entering the co-location environment for the first time is cross-connect pricing. Physical cross-connects between your cage and a venue's presence in the same data center run $250 to $800 per month per connection depending on data center and venue. If you are connecting to five venues in NY4 and one in NASDAQ's Carteret data center, you are paying cross-connect fees to five venues plus potentially a separate cage in Carteret.

These costs accumulate and represent a floor on the monthly infrastructure cost for a fully co-located routing operation. They should be in any honest comparison of co-location versus cloud pricing.

What the Numbers Tell You to Do

The answer is not surprising but it is specific: if your strategy requires sub-500-us tick-to-order latency and you are trading at venues with co-location availability, co-location in the venue's primary data center with kernel bypass networking is the correct architecture. The latency difference is not incremental. It is an order of magnitude.

If your latency requirement is in the milliseconds range and your operational team is small, cloud provides a better cost-adjusted architecture. The 180 us p50 from a properly configured EC2 instance with enhanced networking is fast enough for a large class of strategies.

What does not make sense is taking a strategy that requires 100 us execution and trying to optimize cloud infrastructure to get there. The kernel bypass gap is a fundamental architectural constraint, not a tuning problem.

EurekaLabs

See the Infrastructure Behind These Numbers

Request access to the EurekaLabs platform and run your own latency benchmark against your current infrastructure.