Back to Pulse
EurekaLabs Pulse

Normalizing 12 Feed Formats Without Adding Latency

Priya Venkataraman
Normalizing 12 Feed Formats Without Adding Latency

NASDAQ ITCH uses a binary protocol where each message type is a fixed-width byte sequence with no framing overhead. CME MDP 3.0 uses SBE with a template registry. CBOE PITCH packs order IDs and prices into specific bit positions. ICE's proprietary binary format has a versioned header that requires a schema lookup before you can read the body. FAST compression layers an adaptive variable-length encoding on top of whatever the underlying schema is.

Every feed speaks a different dialect. If you are connecting to five venues, you are handling five distinct decode paths, five different order book semantics, and five different sequencing models. Connecting to fifteen venues means fifteen of all of the above. The normalization problem is not just a protocol translation exercise. It is a low-latency systems problem that gets more complex at each new venue you add.

What Normalization Actually Means

The goal of normalization is to convert heterogeneous feed messages into a single internal representation that the rest of your system can reason about without knowing which venue or protocol produced the data. This internal representation needs to capture: event type (add/modify/cancel/trade), instrument identifier in a canonical form, price in a common precision model, size, side, and a hardware timestamp that reflects when the event was observable in your infrastructure.

What normalization does not mean is that you homogenize away the differences that matter. ITCH includes MPID (market participant identifier) on order adds, which is absent in some other protocols. CBOE PITCH includes a timestamp in most message types at nanosecond precision. These fields are worth preserving as optional extensions on your internal representation rather than discarding them. Your routing and analytics layers may want them.

The Two Architectures and Why One Wins on Latency

There are two broad architectural approaches to feed normalization. The first is a centralized normalizer: all raw feed data flows into a single process that decodes, normalizes, and republishes to consumers. This is operationally simple but adds a hop to the data path. Every latency-sensitive consumer is paying the serialization cost of the central normalizer plus the inter-process or inter-host communication cost.

The second approach is per-consumer normalization: each consumer embeds its own feed handler library and decodes directly. This eliminates the normalization hop at the cost of code duplication and per-feed expertise being spread across multiple teams. The hybrid approach is what most serious trading infrastructure ends up with: a shared feed handler library that each consumer links in-process, combined with a normalized multicast stream for non-latency-sensitive consumers like analytics and surveillance.

We build the EurekaLabs feed handler as a library, not a service. The routing core links it directly and receives decoded events via a lock-free single-producer single-consumer ring buffer. The normalization cost is paid inline with the routing logic, in the same thread, with no serialization or IPC overhead. Our normalized analytics stream runs as a separate subscriber that receives events from the same ring buffer with a slightly higher read pointer.

The Decode Budget Per Protocol

Not all feed protocols cost the same to decode. In our testbed (co-located, Intel Ice Lake, DPDK receive path), we observe the following median decode times per message type:

  • NASDAQ ITCH 5.0 add order: 0.35 us
  • CME MDP 3.0 incremental refresh: 0.48 us
  • CBOE PITCH add order: 0.31 us
  • SBE-encoded binary (generic): 0.40 to 0.55 us depending on template complexity
  • FAST-compressed (any underlying schema): 0.80 to 1.1 us

FAST stands out as the expensive case. The adaptive Fibonacci encoding requires a state machine that cannot be vectorized, and each field decode depends on the output of the previous one. If you are consuming FAST feeds in a latency-sensitive path, this is a real cost. In some cases the better trade-off is to subscribe to the redundant uncompressed version of the same feed where it is available.

Instrument Identifier Normalization

One of the underappreciated costs in multi-venue normalization is instrument identifier translation. Each venue uses its own symbol scheme. NASDAQ uses the stock ticker symbol directly. CME uses a numeric instrument ID from the security definition. ICE uses its own product code. Your internal order book needs one canonical identifier per instrument so that your routing logic can compare order book state across venues for the same underlying.

The naive approach is a hash map lookup on every message. At the volume of a top-of-book feed for the full equity universe (roughly 8,000 to 12,000 active symbols on a busy session), this is acceptable. At NMS data rates for Level 2 depth updates, the hash map lookup adds up. We pre-compute a translation array indexed by the venue's native integer identifier where possible, falling back to hash map only for string-keyed symbols. This reduces the translation cost from roughly 120 ns to under 20 ns for the common case.

Timestamp Harmonization

Each feed delivers timestamps in its own format with its own precision and epoch. ITCH timestamps are nanoseconds since midnight UTC. CME MDP uses a Unix epoch in nanoseconds. Some feeds deliver exchange-side timestamps, others deliver gateway-side timestamps. Harmonizing these into a single timeline requires knowing the timestamp semantics of each feed and applying the right offset.

Our normalization layer preserves both the original exchange/gateway timestamp from the feed and a hardware receive timestamp applied by our NIC at packet arrival. The delta between the two is useful for monitoring feed latency and detecting when a venue's published timestamps drift or lag. If you only record the receive timestamp, you lose the ability to distinguish between a fast exchange and a fast network connection.

Testing the Normalization Layer

A normalized feed handler needs a testing strategy that goes beyond unit tests on individual decode functions. Two things we found necessary: first, a corpus of real captured packets from each feed covering edge cases (crossed markets, trade corrections, out-of-sequence messages), verified against a known-good reference implementation. Second, a replay harness that can inject these packets at realistic rates and measure normalization latency under load.

The normalization layer should be transparent to the routing engine, but its behavior under stress (high message rates, malformed packets, gap detection on UDP feeds) determines system reliability. We treat normalization test coverage as on par with routing logic coverage. A latency regression in the normalization layer caused by a feed schema update is just as costly as a routing regression.

Adding a new feed format is a well-defined engineering task once you have the library architecture and test infrastructure in place. The variable is how different the new protocol's sequencing model is from what you already handle. ITCH-family protocols are easy additions. FAST-based feeds require more effort. Proprietary binary feeds require a reverse-engineering phase if the vendor does not publish a complete specification.

EurekaLabs

See the Infrastructure Behind These Numbers

Request access to the EurekaLabs platform and run your own latency benchmark against your current infrastructure.