the.bay.news

Running the WUIC assistant on a local LLM: Ollama, an MCP server, and a free agentic VS Code

DEV Community
Running the WUIC assistant on a local LLM: Ollama, an MCP server, and a free agentic VS Code
We moved the generative half of our RAG chatbot off a paid cloud API and onto a local model served by Ollama on a GPU box — and exposed the same WUIC knowledge to VS Code through a tiny MCP server, which eventually grew into our own downloadable extension. This is the architecture, the one hard problem (tool-calling), the measured results, and an honest accounting of what a local LLM costs you in quality and latency to save you in money and privacy.

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.