Toward Sustainable On-Device Intelligence: A Survey on Energy-Efficient RAG Systems with Small Language Models — Stanford
Research Square
Toward Sustainable On-Device Intelligence: A Survey on Energy-Efficient RAG Systems with Small Language Models — Stanford
The proliferation of Large Language Models (LLMs) has driven a paradigm shift toward on-device inference, motivated by privacy preservation, latency reduction, and offline capability. Simultaneously, Retrieval-Augmented Generation (RAG) has emerged as the dominant pattern for grounding languag...
0 comments
No comments yet.