Why Routing Systems Fail During High-Volatility Events
Most routing failures during high-volatility periods are not caused by the routing algorithm itself. They are caused by the inputs the algorithm receives. A router that scores venues based on historical fill rates and current quoted spread will perform poorly when those inputs become unreliable, and during rapid market moves, both inputs degrade simultaneously.
Quote stuffing, where a market participant floods an exchange with a large number of orders and cancellations to slow competing systems, causes two problems. First, your feed handler queues build up, and the market state your router sees falls behind the actual market by tens or hundreds of milliseconds. Second, the spread and depth figures in that stale state are misleading: they reflect the order book a moment ago, not now.
Flash events create a related but different problem: the market moves faster than your normal routing decision cycle. If you are checking venue scores every 5 milliseconds under normal conditions, a 200-millisecond flash crash can be mostly over before your routing engine has updated its venue preferences. Orders submitted based on pre-event venue scores may route to venues that are now on the wrong side of a price dislocation.
Detecting Market Stress in Real Time
Before your routing system can adapt, it needs to recognize that conditions have changed. The detection problem is harder than it looks because the signals are noisy and you cannot afford many false positives. Incorrectly classifying normal elevated volatility as a stress event and entering a conservative routing mode wastes execution quality on an ordinary trading day.
The signals we track are: feed lag (how far behind real time is the message sequence number on each venue, measured by comparing expected versus received sequence counts over a rolling 100ms window), quote instability (the rate of add/cancel cycles per price level per second, which spikes during quote stuffing), and spread volatility (the variance of the best-bid/best-ask spread over the last 30 seconds compared to the prior session average). Each of these individually can be elevated during normal activity. All three elevated simultaneously is a reliable stress indicator.
The threshold for triggering stress mode is not a single number. We calibrate it per symbol based on the symbol's baseline volatility profile. A threshold that works for an S&P 500 ETF may be far too aggressive for a mid-cap equity. The calibration needs to run continuously against recent session data, not be set once at system startup.
What Routing Adaptation Actually Looks Like
When the stress detector fires, the routing engine switches from its normal venue-scoring mode to a stress-adapted mode. The changes are specific and measured, not a general fallback to safe behavior:
First, reduce the routing decision window. Under normal conditions, we batch small orders and route them together to optimize for queue position. Under stress, we route immediately without batching. The cost is slightly worse queue position; the benefit is that we are routing against the current market state rather than a state that was current 20 milliseconds ago.
Second, penalize venues with elevated feed lag in the venue score. A venue whose feed is 50ms behind real time is effectively invisible during fast moves. Its quoted prices are stale, and routing to it based on those prices often results in a fill at a price that was valid when your routing decision was made but is now significantly off-market.
Third, increase the aggressiveness threshold for marketable orders. Under normal conditions, we sometimes use limit orders priced aggressively to capture favorable queue position. Under stress, the risk of the market moving through your limit price before the order is processed increases sharply. We shift more volume to market orders or tighter-legged limit orders with a shorter time-in-force.
Venue Instability and Automatic Failover
Venue instability during stress periods is a separate concern from general market stress. Sometimes one venue has connectivity problems while others are functioning normally. If your feed from a normally-preferred venue is producing gaps and your fill acknowledgments from that venue are delayed, continuing to route orders there is counterproductive.
The venue health tracking in our routing engine maintains a health score per venue that updates on three inputs: feed sequence gap rate, order ACK latency (compared to that venue's baseline), and the fill rate of recently routed orders. A venue can pass the feed health check and the routing health check and still fail the fill rate check if orders are being submitted and timing out rather than filling. All three need to be healthy for a venue to maintain its full routing weight.
Automatic failover when a venue's health score drops below a threshold is not the same as cutting off the venue entirely. We reduce its routing weight proportionally rather than removing it from consideration. A venue at 40% health still gets some order flow, which continues to generate data for the health scoring system. Completely removing it creates a cold-start problem when the venue recovers.
What Adaptive Routing Does Not Fix
Being clear about the limits is important. Adaptive routing improves execution quality during stress periods compared to a static routing model, but it cannot overcome latency disadvantages caused by infrastructure. If your feed handler is 100ms behind the market because your connectivity is not co-located, no routing algorithm adaptation compensates for that. The routing engine can only optimize the decisions it makes; it cannot fix the staleness of the inputs it receives.
Adaptive routing also cannot prevent adverse selection during genuine one-sided market moves. When a large macro event causes a sustained price dislocation across all venues simultaneously, routing optimization reduces the magnitude of adverse selection (by preferring venues with better fill rates at that moment), but it cannot eliminate it. The quality of execution during a true market dislocation depends on the quality of your signal to recognize the dislocation before routing, not on the routing engine itself.
This is a distinction worth preserving when evaluating execution infrastructure. The routing layer handles venue-level optimization: where to send an order given a signal. The signal quality and timing are upstream concerns. Excellent routing on a stale signal is still a poor result.