Running the WUIC assistant on a local LLM: Ollama, an MCP server, and a free agentic VS Code
DEV Community
Running the WUIC assistant on a local LLM: Ollama, an MCP server, and a free agentic VS Code
We moved the generative half of our RAG chatbot off a paid cloud API and onto a local model served by Ollama on a GPU box — and exposed the same WUIC knowledge to VS Code through a tiny MCP server, which eventually grew into our own downloadable extension. This is the architecture, the one hard problem (tool-calling), the measured results, and an honest accounting of what a local LLM costs you in quality and latency to save you in money and privacy.
0 comments
No comments yet.