Back to Pulse
EurekaLabs Pulse

Storing Tick Data at Nanosecond Precision: Storage Choices and Tradeoffs

Haruto Yamane
Storing Tick Data at Nanosecond Precision: Storage Choices and Tradeoffs

The Precision Problem

Most databases represent timestamps as 64-bit integers or float values with microsecond precision. That is fine for application logs and event streams where you care about ordering at the millisecond level. It is not fine for market data storage, where two messages from the same venue can arrive within the same microsecond and the ordering between them is meaningful.

Modern exchange feeds, including NASDAQ ITCH and CME MDP 3.0, timestamp events at nanosecond precision using hardware timestamps from exchange-side NIC cards. If your storage layer silently truncates to microseconds, you lose the relative ordering of events that occurred within the same microsecond. For post-trade analysis and compliance reconstruction, this is a correctness problem, not just a precision problem.

kdb+ and Q: The Reference Implementation for Tick Storage

kdb+ has been the default choice for tick data storage at trading firms for over two decades. Its columnar storage model, native temporal types, and tight integration between the storage layer and the Q query language give it genuine advantages for tick data workloads. The nanosecond timestamp type (`n` in Q) handles full precision correctly.

The tradeoff is cost and portability. kdb+ licensing is significant, and the Q language has a steep learning curve that concentrates knowledge in a small group. For firms building new infrastructure in 2025, kdb+ is the right answer if you have Q expertise on the team and are processing tick volumes that justify the investment. It is the wrong answer if you are evaluating storage options without that expertise already present.

TimescaleDB: PostgreSQL Semantics at Tick Volumes

TimescaleDB is a PostgreSQL extension that adds time-series hypertable partitioning and compression. Its appeal is that it uses standard SQL, runs on commodity hardware, and integrates with the existing PostgreSQL ecosystem. The nanosecond timestamp issue is manageable because PostgreSQL's timestamp type stores up to microsecond precision natively; true nanosecond storage requires using a bigint column and handling the conversion explicitly.

In our lab tests against a simulated ITCH feed replay at roughly 500,000 events per second, TimescaleDB with compression enabled and appropriately sized chunk intervals handled write throughput without observable lag at that rate. Query performance on historical symbol slices within a time window was fast. Full-day scans across a large symbol universe were the weak point, particularly when the query requires sorted output by event time across multiple hypertable chunks.

TimescaleDB is a reasonable choice for operations teams who are more comfortable with PostgreSQL than columnar systems, and for workloads where query flexibility matters more than raw read performance at the extremes.

ClickHouse: Columnar Performance for Analytical Queries

ClickHouse stores data in a columnar format optimized for aggregation queries. For tick data analytics, "how did spread behave for symbol X between 9:30 and 10:00 over the last 30 trading days" is a query ClickHouse handles extremely well. It will outperform TimescaleDB on analytical aggregations once the dataset is large enough that full scans of a column are significant.

Where ClickHouse is weaker: real-time write latency is not its design priority. It buffers writes internally and flushes periodically, which means the most recent data may not be immediately available for query. For systems where current-session tick data needs to be queryable within seconds of arrival, you typically run a dual-write strategy: current data goes to a fast in-memory store or time-series DB, historical data migrates to ClickHouse for long-horizon analytics.

ClickHouse also handles nanosecond precision via a Datetime64(9) type, which is a first-class type in recent versions, making it cleaner than the PostgreSQL workaround.

Custom Binary Formats: When You Need Full Control

Some teams write their own binary format for tick data, typically a memory-mapped file per symbol per day with a fixed-length record structure. The record contains: nanosecond timestamp as a raw uint64, event type byte, price as a fixed-point integer (to avoid floating-point representation issues), quantity, and any protocol-specific fields. The format is write-once, read-many. Reads are sequential for replay or random-access indexed by timestamp for point queries.

The advantages: zero serialization overhead, predictable memory layout, trivially parallelizable reads across multiple files, and full nanosecond precision with no third-party storage dependency. The disadvantages: you own everything. Schema changes require migration tooling. Ad hoc queries require writing custom readers. Backup and recovery is your problem.

For high-frequency trading teams where the storage layer is on the hot path of real-time replay (feeding a signal engine, not just storing for later analysis), custom binary formats are often the right answer despite the maintenance burden. The overhead of a storage abstraction layer is measurable when you replay at microsecond granularity.

What the Choice Actually Depends On

The right storage choice is not universal. It depends on your dominant query pattern. If most of your access is time-range slices of a single symbol for intraday replay: kdb+ or custom binary. If most of your access is multi-symbol, multi-day aggregations for strategy research: ClickHouse. If your team has PostgreSQL expertise and your tick volumes are in the low hundreds of thousands per second: TimescaleDB is a pragmatic choice that reduces operational complexity.

Where we would push back on common assumptions: the idea that you need to pick one system for all workloads. In practice, a hot-path custom binary format for current-session replay and a ClickHouse cluster for historical analytics handles most needs without forcing you to accept the weaknesses of either system everywhere. The migration between them is straightforward given that you own the record format on both ends.

A Note on Compression

Tick data compresses well because price increments are small and the data is ordered by time. Delta encoding on price and timestamp fields before compression is standard practice. For reference: uncompressed ITCH feed data for a full-day equities session on a major venue runs roughly 80-120 GB. With delta encoding and LZ4 compression, this typically reduces to 15-25 GB depending on the session's volatility and message rate. Storage costs at this scale are not a primary concern, but I/O bandwidth during full-day replay of multiple symbols in parallel can be, particularly on spinning disks.

EurekaLabs

See the Infrastructure Behind These Numbers

Request access to the EurekaLabs platform and run your own latency benchmark against your current infrastructure.