logoalt Hacker News

samtp • yesterday at 6:09 PM • 3 replies • view on HN

After reading the post, still don't fully understand when you would use CF Queues vs K2. Can you help to elaborate a bit more?


Replies

pcthrowaway • yesterday at 11:03 PM

My understanding: Distributed queues are generally good for when you have multiple workers processing chunks of work. Each queue item usually needs to be processed to completion exactly one time, so the queue provides the mechanism for the workers to coordinate state of each item at the item level (which allows for time-outs and re-tries if, say, a worker dies during processing, like if you were using spot instances for your worker pool)

Event streams are for multiple consumers, and can offer different guarantees. As far as I can tell, K2 is designed to ensure all consumers receive all events at least once (it's unclear to me whether this means they're continuously storing all events from the stream origin, or if older events age out at some point, or are dropped when they've been consumed by all known consumers).

Other types of guarantees with streams might be "at-least-once", "at-most-once", and "exactly-once" delivery, for different needs. Redis streams used to be at-least-once but it looks like they support all 3 use cases now. Some relational DBMSes also have the option to replicate by streaming their transaction logs to all servers in the cluster so each node maintains its own understanding of the database state (though stale reads can also occur in some/all? DBMSes that replicate this way, when a server is queried before receiving an update)

necubi • yesterday at 7:45 PM

Sure! There are definitely some overlapping use cases, and we've seen folks using/abusing queues for use cases that are more appropriate to something like k2.

Queues are great when you have a unit of work that needs to be completed, retried, and tracked individually. For example, a shop might need to call a payment processor API that can fail or timeout, and retry it until it succeeds, while polling on the frontend for the state of that particular message. With a queue, you can insert a message tracking that payment, and have a queue processor that keeps getting sent it until it succeeds or has failed too many times.

In a queue each item is its own thing that's important to someone, and queues give you APIs to interact with that particular item.

K2 is for moving large volumes of data around. Pricing is per GB, not per message. Records are produced and consumed in bulk, and what matters is that all records are processed, but no one is querying the state of a particular record. K2 also supports multiple consumers for the same record, and long term retention. For example, all of your applications may emit events when things happen, and those events need to be read by an alerting system, a system that durably stores them, and a system that uses them to build ML features.

➕ show 1 reply
hasyimibhar • yesterday at 8:52 PM

Queues are for actions (do this), K2/event streams are for events (this happened).