Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
This is why I think llms are a killer app for Linux desktop. They’ve been trained on Linux very hard, and it cleanly solves the “how do I make it do $thing” problem since everything is open and the llm can manipulate it. For example: Sound not working? Just tell the llm.
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
> it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo.
Most llm could do it. Claude went from firmware thread -> rtos scheduler -> mcu reference manual -> hardware controller register interface -> vendor sdk -> problem identification and the solution to it in a matter of 30 minutes. Linux could be even easier since it is so well trained on.
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
3.8 Flash is my daily driver and produces pretty excellent results all round.
I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
My theory is that Gemini 3.8 Flash was supposed to be Gemini Pro, but by the time it was ready for release, Google was embarrassed at how far behind their "pro" model was, so they just named it Flash. It's a good model, but don't let the name trick you.
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
So Gemini Pro (I use agy, because no other harness can be used for subscription plan) was supposed to get Gemini 3.5 Pro, at some point (it was planned, right?), but instead Argon arrives and I am pretty sure it will only be in Ultra sub and don't think 3.5 Pro is coming anymore.
> I was trying to use rocm with llama.cpp
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
``` void* mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) { if (!real_mmap) real_mmap = dlsym(RTLD_NEXT, "mmap"); ```
hope it's not run by multiple threads and dlsym is not allocating.
I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it's not as capable but it's fast and almost free (sufficient quota with pro account). The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
That's cool, but the real question is, why aren't you using Hipfire[1] or HaloPFX[2]? Both are far superior to llama.cpp in terms of performance, for Strix Halo.
I would like to use my Google AI Pro subscription included with Drive but last time I checked, their terms were not just ambiguous and confusing wrt training and ZDR, they were contradictory. Until they have that fixed and properly communicated, I'll stick to other vendors.
Adding my anecdote, because it amused me: I finished wiring up the compute/sensor box for my robot, ssh'd in and told agy "I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output". It did all the local config for the lidar, setup docker and the ROS2 environment, then told me "open up this url in foxglove" and sure enough everything worked. Whole thing used up 6% of my weekly limit.
Please tell me you published your findings even as an issue on the llama.cpp GitHub
if you haven't tried Qwen3.8-Flash-Next with halogen, you're missing out: https://github.com/peonist-ai/halogen-flash-server#the-host-...
lol I've hit the hipStreamCreate problem!
[flagged]
[flagged]
I had a similar but less impressive experience recently with Muse Spark 1.3.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.