Hacker News
Show HN: PicoMQ – Durable Streams over HTTP, on object storage
S3Stream is the stream storage primitive, used in AutoMQ, shipped as a Rust library. Coordination is a command log in Postgres.
gpsmsn
|next
[-]
Initially, I started to compare PicoMQ to https://github.com/s2-streamstore/s2. If you're familiar with them, how would you compare PicoMQ to S2?
adesh_nalpet
|root
|parent
[-]
SlateDB is great, but it's not purpose-built for streaming workloads. With PicoMQ, the goal is to really do one thing well. It would have been far easier to just extend SlateDB, but instead, the S3Stream engine is built from the ground up with AutoMQ's core primitives.
Not to mention, Pico supports multi-node/cluster deployments and a bunch of other features (no feature gating).
Jonovono
|next
|previous
[-]
Have you looked into the semantics of their like StreamDB stuff ? Pretty interesting
Also there is this that was built on top of it. Sounds somewhat similar goal as yours but I could be off. Need to explore more https://ursula.tonbo.io/
adesh_nalpet
|root
|parent
[-]
I did, in fact, initially write PicoMQ with OpenRaft, similar to Ursala, but I really wanted the operational complexity to be minimal and the nodes to be stateless (at least for most use cases), Pico uses SQL database as a metadata command log, inspired by RisingWave.
But Ursala would, without a doubt, have better durability ACK latency, as it wouldn’t have to wait for an ACK from S3. That said, I do plan on extending disk/EBS-staged WAL to hit similar low-latency durability ACK numbers. But again, most use cases don’t need single-digit-millisecond latency for ACKs.
Jonovono
|root
|parent
[-]
adesh_nalpet
|root
|parent
[-]
KaiserPro
|next
|previous
[-]
This sounds like a kafka-like streaming system, but backed onto s3-objects?
Doesn't this mean that write performance is going to be bad?
adesh_nalpet
|root
|parent
|next
[-]
And it's also fair to question write performance, since it's backed by object storage. The optimization is primarily from the shared WAL across streams, server-side batching, and client-side in-memory pipelining, especially with HTTP/2, without as much connection pool overhead.
In practice, you can go to the extent of achieving up to 100 MiB/s throughput per stream. Considering how granular streams can be, you'd rarely need as much. The latency for a durability ACK is, however, the price to pay, which is going to be ~250 ms, or lower with S3 Express, which I'd say covers most real-time use-cases. The design itself is easy enough to extend to a disk-staged WAL for single-digit durability ACK latency.
atombender
|root
|parent
[-]
Shakahs
|root
|parent
|previous
[-]
Regular S3 has write latency 100-150ms, which might be fine depending on your workload anyway.
soleveloper
|next
|previous
[-]
So - in theory - something like a massive chat client, discord like, can be implemented via this solution? And what would be the pricing of such a solution. Cheap-serverless-discord
adesh_nalpet
|root
|parent
[-]
Exactly, streams can essentially be rooms, and since the ordering is preserved, a Discord-like application is a strong use case. I’m even considering building one using PicoMQ as an example showcase.
The pricing is going to be dirt cheap, and the best part is how easy it is to scale up vertically or add nodes. For some raw numbers, assuming 1M messages/day, 200B per message, ~6GB/month, all-inclusive, it would be $30 to $150 a month, and storage would be the cheapest part.
soleveloper
|root
|parent
[-]
Is this back of envelope pricing include the traffic/bandwidth of the readers?
That could easily be 10-100x of number of messages.
This solution together with a cheap/free caching layer (especially for non members/writers) could be amazing.
BTW, a classic example would be an hn mirror. ;)
adesh_nalpet
|root
|parent
[-]
But as we speak, I am in the process of running open-benchmarks on AWS with cost attribution. I'll be posting them in the docs with transparency soon enough.
Read and write through cache already exists today, which would work in favour of both cost, latency, and consumer reads fanout.
I like the HN Mirror over a Discord-like app for simplicity. No doubt that's where my weekend is going!