logoalt Hacker News

dudeinhawaiiyesterday at 6:22 PM2 repliesview on HN

I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but they never bothered building or running. I know a lot of this can be fixed with workflows but it still feels like a failing.

GPT-5.6 or Claude models haven't delivered to me non-running code in ages.

Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output.

I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome.

As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided.

I don't think it's a major selling point when every model can do it well and reasonably fast.

That said, eagerly awaiting "pro" and improvements to antigravity.


Replies

WarmWashyesterday at 9:41 PM

It depends on if you are relying on all single shot tasks or are willing to iterate. Flash is quick and can make dumb mistakes, but it also can fix them quickly.

I've gotten good results with it, but it definitely is more hands on.

boinkboink78912yesterday at 6:51 PM

Try this one, it's a step jump in coding capabilities for me over 3.6.