logoalt Hacker News

rwzyesterday at 8:10 PM2 repliesview on HN

0.5t/s is still too slow even for email. For a moderately large inquiry (1MTok output, let's ignore the 4k context window limitation for now) it'll take the model around 23 days or uninterrupted execution to answer a single email.

Real world inquiries are gonna be much slower of course, but this setup is still too slow to do anything meaningfully useful I think.


Replies

0cf8612b2e1etoday at 12:05 AM

Sure, you cannot do 1M, but there are plenty of useful questions you could ask that are far more modest. Simple Q+A, look at this function, how would you design X? All of those could have few paragraphs of outputs that would finish within a day.

springtimesunyesterday at 9:23 PM

I’m sure it’s possible, but I really struggle to think of an example that would result in a 1m token output.

show 1 reply