logoalt Hacker News

WebLLM: high-performance in-browser LLM inference engine

98 pointsby saikatsgyesterday at 2:02 PM17 commentsview on HN

Comments

TekMolyesterday at 3:12 PM

This seems to be the demo:

https://chat.webllm.ai/

I am getting:

    WebGPUNotAvailableError: WebGPU is not supported in
    your current environment, but it is necessary to
    run the WebLLM engine.
On both, FireFox and Chromium on Linux.
show 2 replies
mandeepjyesterday at 9:37 PM

It is kinda obvious, but maybe that's why it's not stated anywhere: each browser session will result in a download of 500 MB to ~1 GB, depending on your model selection. So, it's better to add a disclaimer if you end up using WebLLM in a customer-facing site.

MarioManyesterday at 7:25 PM

I really enjoy this engine. I’ve used it for personal projects, but it hasn’t been updated since Gemma 2. I suggest using Transformers.js instead these days.

refulgentisyesterday at 2:44 PM

Project is de facto dead, used it for many years and had to rip it out 6 months ago, don't waste your time.

show 1 reply
init0yesterday at 6:36 PM

You might like webml-kit https://npm.im/webml-kit

conceptmeyesterday at 5:47 PM

Bake me a cake

responds with

> Error: Cannot initialize runtime because of requested maxStorageBuffersPerShaderStage exceeds limit. requested=10, limit=9.

adastra22yesterday at 6:21 PM

A WebX technology that actually involves browsers!

paidxtoday at 1:12 AM

[flagged]

nnevatieyesterday at 3:59 PM

[flagged]

asynczeyesterday at 4:34 PM

[dead]

ycCantCodeyesterday at 5:30 PM

[dead]