logoalt Hacker News

GigaToken: ~1000x faster Language model tokenization

230 pointsby syrusakbarytoday at 5:20 PM45 commentsview on HN

Comments

maxdotoday at 5:57 PM

Interesting :

Q: Did you just way over-optimize for a specific CPU and tokenizer? How is it so fast? No, I way over-optimized for every combination of these! The results are very consistent across CPUs (modern x86 and ARM), and across specific tokenizers.

The major improvements are in optimizing heavily an implementation that usually is outsourced to a Regex engine (pretokenization) using SIMD, minimizing branching and other tricks, as well as heavily optimizing caching of pretoken mappings (if a word has been seen before, look it up its encoded tokens efficiently). Caching is a very hard problem in this domain since the cache grows very quickly, and pretoken distributions are very long-tailed.

Finally, interactions with Python are minimized, and threads have minimal interactions with each other.

onlyrealcuzzotoday at 6:31 PM

This is awesome, but tokenization is typically <0.1% of total inference time.

Presumably there's a host of applications that just need to tokenize, though, and this would be great for those!

show 5 replies
luciana1utoday at 8:22 PM

engineering effort to make something 1000x faster that accounts for 0.1% of total runtime is the most software developer thing imaginable

show 3 replies
0xnyntoday at 6:28 PM

I had to stare at that chart for a minute just to let the numbers sink in. It's genuinely mind-bending, incredible ship OP

chocratestoday at 8:25 PM

Practically I would need to wait for hugging face models to adopt this? My harness tokenizer is just an estimate since the model tokenizes on my api calls?

swiftcodertoday at 7:52 PM

So the question becomes, how many other parts of the inference pipeline have left 1000x optimization opportunities lying on the table?

show 3 replies
fwiptoday at 5:57 PM

What sort of setups do people have that are bounded by the speed of the tokenizer?

show 6 replies
sashank_1509today at 6:05 PM

This is really cool, great work!

anonymousmoostoday at 7:12 PM

Quality software here.

dmezzettitoday at 6:54 PM

Very interesting project! Are there benchmarks for the "compatibility mode" or are all the numbers for the Gigatoken API?

show 2 replies
zerolinestoday at 6:47 PM

wow, best release all week.

semiinfinitelytoday at 6:41 PM

quite excellent software

vmware508today at 7:56 PM

We should just rewrite everything in Rust, especially bloated Python code, and the world would be a better place. ;) Disclosure: I'm a Rust advocate!

show 4 replies