There's also token RTT on top of network latency.. but what if you had a model running at 10k tps (like taalas' llama3b-8