logoalt Hacker News

misiti3780yesterday at 4:16 PM8 repliesview on HN

is anyone doing this ?


Replies

LarsDu88yesterday at 4:20 PM

The closest example I've seen is ChatJimmy: https://chatjimmy.ai/ a prototype from Taalas running Llama 8B

Scaling this up to 2.8 Trillion (350X increase), will certainly be challenging.

If I was younger and had the right background, I'd love to dive into attempting somethign like this

show 1 reply
alach11yesterday at 11:50 PM

News just broke today that Google is planning on doing this: https://news.ycombinator.com/item?id=48986351

gopalvyesterday at 4:44 PM

Taalas HC1 is the closest thing to this.

Last I saw they posted Deepseek R1 numbers in Feb of this year.

The challenge is rolling out a new one every 7-8 weeks as the weights change & cheap enough for a hyper scaler to afford to buy one and save enough on power over the next 8 weeks as a payoff.

carterschonwaldyesterday at 4:20 PM

not sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out

535188B17C93743yesterday at 4:51 PM

Yeah, I've heard of Taalas doing it. Not sure of others but I'm sure lots of companies are considering it, especially as we start to hit points of depreciating returns in training.

eckryesterday at 4:20 PM

There was a startup that did this for Llama 3, I forgot their name. Etched is also doing some similar things I believe.

robgoughyesterday at 4:22 PM

obligatory link to https://chatjimmy.ai

hnfongyesterday at 4:35 PM

It sounds vaguely similar to what Cerebras.ai is doing