Hacker News
Show HN: Parseable, an open observability datalake, handles 100M time-series/min
nwmcsween
|next
[-]
parmesant
|root
|parent
[-]
a) there's no per-series inverted index and labels are parquet columns so memory is not bounded by cardinality
b) data lives on much cheaper object storage (parseable gives an option to cache data locally to remove io bound latency)
c) columnar store helps with faster data scanning by aggressively pruning and filtering data out
yashdotrv
|next
|previous
[-]
This is Yash, founding team at Parseable (https://github.com/parseablehq).
We've built an open source observability data lake using Rust, that handles high-cardinality data at around 100M time series in production (https://www.parseable.com/blog/how-parseable-handles-100-mil...)
Our architecture is built around columnar design, and we use Apache Arrow for in-memory columnar processing and Apache Parquet for durable columnar storage on S3-compatible object storage. In Parseable, every labels stay as columns in the data instead of becoming a large long-lived per-series index like many TSDBs.
Also, one thing we’ve been thinking about a lot is how observability changes as agents become part of day-to-day engineering workflows. They're not just another service, they produce traces, tool calls, prompts, intermediate decisions, errors, costs, and sometimes sensitive business context.
Observing them matters just as much as observing any other system. But it is equally important to decide where that telemetry data should reside. Our view is that teams should be able to keep these observability data close to them: in their own object storage, under their own retention, access, and compliance controls.
codegeek
|root
|parent
|next
[-]
goldeneye13_
|root
|parent
|next
|previous
[-]
parmesant
|root
|parent
|next
[-]
nikhil4usinha
|root
|parent
|previous
[-]
The reason we think it scales differently - labels are just columns in Parquet, so there is no per series index that grows with cardinality. In that deployment one label alone has ~2.5M distinct values among 500+ labels, which would be painful for an index based TSDB but here is just a high cardinality column. What drives cost for us is ingestion rate (data points/s) and how much data a query has to scan for a particular time range not series count. Ingest scales horizontally by adding ingestors, and queries prune by time partition and column stats.
A billion series benchmark is on our list, and we'll publish the numbers when we run it.
msandford
|root
|parent
|next
|previous
[-]
nikhil4usinha
|root
|parent
[-]
gustavohoa
|root
|parent
|previous
[-]
parmesant
|root
|parent
[-]
5x Queriers, each with 64 vcpu 192 GB
The current utilization sits comfortably at 10-15 vcpu and 20-30 GB memory for the ingestors 20-40 vcpu and 40-60 GB memory for the queriers
Ample of headroom for transient spikes and planned near-future growth
simonw
|next
|previous
[-]
I found this out because I set Codex the task of running this locally agains another of my apps and it worked around the limitation by running this proxy: https://gist.github.com/simonw/b0e61a0aa8e3f7d30f27ce2f747c9...
... but it turns out my stack can emit JSON just fine, so I switched to that instead. Here's me TIL write-up of getting Parseable running locally https://til.simonwillison.net/datasette/datasette-parseable-...
usernametaken29
|previous
[-]
nylonstrung
|root
|parent
|next
[-]
What could I do with this that I couldn't achieve with iceberg-rs + DataFusion + parquet/vortex
They repeatedly talk about "80% size reduction with compression. Isn't that essentially just the default parquet compression ratio?
What exactly is unique here