logoalt Hacker News

Building a RAG pipeline for semantic code search

31 points • by saikatsg • today at 5:51 PM • 7 comments • view on HN

Comments

keeda • today at 8:23 PM

I think something like this would be key to improving the quality of coding agents. A very common issue (maybe the biggest one) observed by many people is that agents often produce a lot of duplicate and redundant code; multiple abstractions, methods, classes, data structures, etc. serving minor variations of the same purpose... sometimes within the same file!

My theory is that this is due to a kind of "tunnel vision" these models have as they execute on a given task, because engineers new to a company do the same thing until they learn the "lay of the land" and figure out that similar problems have been solved elsewhere.

In a past job my team owned the internal multi-repo codesearch tool, which was by far the most popular internal tool, and later another team added a similar semantic search capability. This was very exciting, but I left before I could see how well it worked out in real-life.

Like, you'd do a keyword search and explore if you need some major piece of functionality that would require significant work, or whenever you encounter an abstraction whose code does not exist in your repo and you want to learn more about it. But when you're in the flow and inventing smaller abstractions, like a class or utility method, you don't necessarily think to search for it. Worse, even if you did, you could not search for it effectively because something similar may exist with slightly different naming or terminology or a typo that a keyword search would miss. Predictably, at scale you ended up with a dozen different implementations doing the same thing.

Now however, you could automate this with agents. I suspect these days simply prompting an agent to look for any relevant code to reuse would actually work pretty well. But they would need to store the entire codebase in their context window (if it fits at all) to refer to it all the time, which would burn a ton of tokens AND reduce performance due to a heavily polluted context. Instead, something like this would be invaluable to provide as a tool / MCP to the agent so that it could locate relevant, reusable code during its planning phase. (Or maybe a post-codegen linter-like check, which IIRC some people have tried, but why fix when you could prevent?)

I'm not sure if the exact method in TFA would work best, though; maybe a pipeline that generates comments/docs for each unit of functionality and then semantically indexes those, rather than structure-aware chunks of code itself?

➕ show 1 reply
duhhhhh1212 • today at 6:58 PM

https://www.pangram.com/history/c901e80e-9cb7-46e4-bf10-7348...

I don't want to say don't waste your time since the first half is human written. Questions for the authors: did y'all just get tired of writing and said "fuck it let's have the LLM finish the rest"? Or did one of you use LLM to write the last half and the other used their own words?

➕ show 1 reply
simianwords • today at 7:15 PM

Here we go again, the industry largely gave on up RAG. In fact I have hardly seen any case where grep doesn't work as well as RAG.

➕ show 1 reply