logoalt Hacker News

ak_tyesterday at 11:23 PM0 repliesview on HN

You don't have to send every single request twice, just the ones that are haven't returned in time. Wait until some threshold, such as your p95 latency, and send your backup request after that. Return whichever request comes back first, and it should cut your tail latency without doubling your cost, since it only duplicates the small % of requests at the tail.

Google calls this a 'hedged request': https://cacm.acm.org/research/the-tail-at-scale/