logoalt Hacker News

taylorfinley • yesterday at 8:16 PM • 25 replies • view on HN

Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...


Replies

spankalee • yesterday at 8:23 PM

3.8 Flash is just quite good, and so is the Antigravity harness.

I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

➕ show 2 replies
plasticchris • today at 12:56 AM

This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.

➕ show 10 replies
IndeanCondor • yesterday at 8:23 PM

Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.

➕ show 1 reply
amanguliani • yesterday at 8:41 PM

Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me

➕ show 1 reply
gottorf • yesterday at 8:44 PM

My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.

➕ show 3 replies
mirmor23 • today at 1:30 AM

> it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.

Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.

yegle • yesterday at 8:45 PM

For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.

And I saw it do this twice, once for Android 14 and once for Android 16.

I think this is just within 3.8 flash's capabilities.

➕ show 1 reply
danpalmer • yesterday at 11:00 PM

3.8 Flash is my daily driver and produces pretty excellent results all round.

➕ show 1 reply
illwrks • yesterday at 10:14 PM

I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.

KMnO4 • today at 12:55 PM

My theory is that Gemini 3.8 Flash was supposed to be Gemini Pro, but by the time it was ready for release, Google was embarrassed at how far behind their "pro" model was, so they just named it Flash. It's a good model, but don't let the name trick you.

➕ show 1 reply
mapontosevenths • yesterday at 8:24 PM

Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.

➕ show 1 reply
crossroadsguy • today at 7:36 AM

So Gemini Pro (I use agy, because no other harness can be used for subscription plan) was supposed to get Gemini 3.5 Pro, at some point (it was planned, right?), but instead Argon arrives and I am pretty sure it will only be in Ultra sub and don't think 3.5 Pro is coming anymore.

➕ show 1 reply
Grimburger • yesterday at 10:38 PM

> I was trying to use rocm with llama.cpp

completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.

once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B

I get the feeling the situation is only going to improve longer term so might be a good time to just do it

➕ show 3 replies
phmx • today at 3:01 PM

``` void* mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) { if (!real_mmap) real_mmap = dlsym(RTLD_NEXT, "mmap"); ```

hope it's not run by multiple threads and dlsym is not allocating.

gcy • yesterday at 10:42 PM

I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it's not as capable but it's fast and almost free (sufficient quota with pro account). The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.

d3Xt3r • today at 3:36 AM

That's cool, but the real question is, why aren't you using Hipfire[1] or HaloPFX[2]? Both are far superior to llama.cpp in terms of performance, for Strix Halo.

[1] https://github.com/warpfront/hipfire

[2] https://github.com/julianmb/halofpx

vinzenzu • today at 7:29 AM

I would like to use my Google AI Pro subscription included with Drive but last time I checked, their terms were not just ambiguous and confusing wrt training and ZDR, they were contradictory. Until they have that fixed and properly communicated, I'll stick to other vendors.

martythemaniak • yesterday at 10:37 PM

Adding my anecdote, because it amused me: I finished wiring up the compute/sensor box for my robot, ssh'd in and told agy "I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output". It did all the local config for the lidar, setup docker and the ROS2 environment, then told me "open up this url in foxglove" and sure enough everything worked. Whole thing used up 6% of my weekly limit.

alightsoul • yesterday at 8:26 PM

Please tell me you published your findings even as an issue on the llama.cpp GitHub

➕ show 4 replies
lardo • yesterday at 10:50 PM

Did rocm provide any benefit over vulkan?

➕ show 1 reply
cyanydeez • yesterday at 10:46 PM

if you haven't tried Qwen3.8-Flash-Next with halogen, you're missing out: https://github.com/peonist-ai/halogen-flash-server#the-host-...

➕ show 1 reply
iknowstuff • yesterday at 11:40 PM

lol I've hit the hipStreamCreate problem!

qrify_app • today at 4:14 AM

[flagged]

la6479 • yesterday at 11:50 PM

[flagged]

bel8 • yesterday at 8:27 PM

I had a similar but less impressive experience recently with Muse Spark 1.3.

Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.

It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.

➕ show 1 reply