Posts tagged: Deployment
What actually holds up in production — retrieval, cost, deployment, and the tradeoffs nobody puts in the sales deck. Written from client work and our own builds.
Open WebUI, AnythingLLM, or LibreChat: the UI question is a retrieval question
All three will give you a clean chat window over your own model in an afternoon, which is why every comparison picks a different winner. The thing that decides it in a year isn't in the feature list — it's who owns the retrieval pipeline.
Read the post →Ollama, vLLM, or llama.cpp: picking a server you won't have to replace
The three are not competing to be the best inference server. They are three different bets on how many people will use the thing at once — and that's the only question that decides which one you want.
Read the post →One post a month, no pitching
Practical notes on private and local LLM systems — hardware, retrieval and costs — in English and Vietnamese. Unsubscribe any time.
Have a workflow you're trying to figure out?
Book a 30-minute discovery call. We'll map it and tell you straight whether it belongs on your hardware or on a cloud API.