I believe embedding-based RAG, everybody is using, will end. As chips advance, you would use a big llm instead of word embedding for retrieval. It's much more accurate and extensive covering every topic.
Still need ~2 years to be replaced.
How would you use a big LLM for retrieval?
How would you use a big LLM for retrieval?