Posts tagged: RAG
What actually holds up in production — retrieval, cost, deployment, and the tradeoffs nobody puts in the sales deck. Written from client work and our own builds.
Open WebUI, AnythingLLM, or LibreChat: the UI question is a retrieval question
All three will give you a clean chat window over your own model in an afternoon, which is why every comparison picks a different winner. The thing that decides it in a year isn't in the feature list — it's who owns the retrieval pipeline.
Read the post →Long context didn't kill RAG — it changed what RAG is for
Every time a model ships a bigger context window, someone declares retrieval obsolete. Then they try to put a real corpus in the prompt and rediscover, at $1.50 a question, why retrieval exists.
Read the post →Vector search alone is not a retrieval system
The demo works and then someone searches for an invoice number. What separates a RAG prototype from something people trust is almost entirely in the retrieval layer.
Read the post →One post a month, no pitching
Practical notes on private and local LLM systems — hardware, retrieval and costs — in English and Vietnamese. Unsubscribe any time.
Have a workflow you're trying to figure out?
Book a 30-minute discovery call. We'll map it and tell you straight whether it belongs on your hardware or on a cloud API.