logoalt Hacker News

TeMPOraL • yesterday at 5:48 PM • 7 replies • view on HN

Jev is one.

Diffusion transformers are not "easy" but underfunded.

Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware - opens up so many possibilities I'm probably unable to imagine half of them.

E.g. Imagine spellcheck/predictive text (or code autocomplete) where the model is able to process a whole paragraph + surrounding application/system context in between keystrokes. Or an OS being able to reliably guess what you're doing in real-time, in between your UI interactions, and offer actually helpful contextual reactions.

Or imagine finally funding some decent studies into exploring the models as computational artifacts - studying their latent spaces, how they form and how they model reality internally.

Or imagine automated sliding doors that don't suck.

--

[0] - Or anything substantially better than BERT-level models used in Jev or that demo from the company doing inference ASICs, that has a chatbot online that does 14 kilotokens per second.


Replies

mdp2021 • yesterday at 7:42 PM

There are "low hanging fruits" - easier to achieve goals -, and there are super-fruits, milestone-fruits.

Among the most important ones:

-- the long-known Problem of Transparency, applied to the apparent emergent intelligence in NNs. Why does it happen - in detail?

-- then, a Theory of Apparent Intelligence through NNs. Transforming the results achieved into a Science. Which allows to do what we are doing - but in a lean and targeted way.

-- then, a General Theory of Intelligence, that includes the above to go beyond current architectures and get those features of Intelligence we expect and still not have.

The long-term direction we got into must lead to this.

(You note a ponderant detail of the above when you note the importance of explaining the emergence of a World Model from a Language Model.)

➕ show 2 replies
blurbleblurble • yesterday at 6:11 PM

Diffusion models combined with these new looping techniques are gonna change the whole conversation about efficiency. Imagine control net but in one or more conceptual latent spaces.

But also harnesses and more generally new insights on "the control flow problem" could end up squeezing a ton of performance out of small models.

flipping_beacon • yesterday at 6:06 PM

Definitely agree with edge computation, although inference extensively researched and funded if SOTA LLMs hit a dead end tomorrow,there is still a lot to explore and research in inference and edge computation

seizethecheese • yesterday at 7:32 PM

I commend you for actually answering, independent of what I think of the answers.

➕ show 1 reply
Amekedl • yesterday at 6:44 PM

yeah your reply, nobody can predict the future.

Enough stuff can happen, software use itself might change, and that could really cause anything. "What will we do with all the gpus" might become a question if for a magnitude of tech and reasons leaked-opus-9 runs on a macbook m6 or 7

aeve890 • yesterday at 6:04 PM

>Random one in terms of applications: getting GPT-4-level[0] LLMs to operate at hundreds of tokens per second on edge hardware

That's low hanging for you?

➕ show 7 replies
alightsoul • yesterday at 7:41 PM

[dead]