logoalt Hacker News

Art9681last Monday at 12:15 AM4 repliesview on HN

The absolute best way to prove this works is by releasing a model that was fine-tuned with this method and then showing benchmarks depicting the improvement delta between the base model and the fine tuned one.

The work is not done. Then release it to the masses and wait a few days for the actual real world anecdotes.

Until then, this is noise.


Replies

SilenNlast Monday at 12:22 AM

Valid criticism. Happy to answer any qs. We're still working on solidfying results.

show 1 reply
irishcoffeelast Monday at 12:20 AM

Benchmarks are the ultimate consolidation of halnons razor.

teravorlast Monday at 12:21 AM

[flagged]

show 1 reply
Reubendlast Monday at 12:19 AM

Yeah, this is just slop. No benchmarks, no concrete case studies, just some vibecoded "platform" to finetune models on your own traces.

Which is an idea that has some value, but also some weaknesses. And this implementation of it isn't forthcoming with that concept. You have to really dig in to understand what they're even talking about.

show 1 reply