Open source · Building in public
local-rag-companion
Self-hostable RAG for companies that cannot send data to a third-party API.
A retrieval system designed for fintech, healthcare, insurance and government, where data residency and audit requirements rule out a cloud LLM provider. Local embedding models, a local LLM, and an OpenAI-compatible surface so existing tooling keeps working after changing one base URL.
Read this before anything else
v0.1 runs, and some of it is verified. The Ollama path has been taken end to end against a real server: real embeddings, real generation, and all seven eval metrics passing. CI runs the suite on three Python versions and against a live Postgres and Qdrant on every push. What has not happened is the full Docker stack coming up, and several adapters have still only met a mocked transport.
Whatever is true today lives in docs/roadmap.md, including a per-backend table of what has been run against a real service and what has only seen a mock. This page deliberately does not keep its own copy of that.
MIT licensed. Python, FastAPI, Postgres. No tagged release, and there will not be one until it has run somewhere real.
search and delete take a tenant id. One place to check for a leak, not a rule spread across every caller.Why anyone would build this again.
Security teams at banks, insurers, hospitals and government departments have blocked cloud LLM providers outright. Not throttled, blocked. The usual answer they get back is "then do not use RAG", which costs those teams the one thing RAG is genuinely good at: answering questions from internal documents nobody has time to read.
Running a model locally is the easy half of that problem and it is mostly solved already. The hard part is proof. A system that passes a compliance review has to show which documents were retrieved, for which tenant, feeding which answer, and that the record was not quietly edited afterwards. Most RAG stacks treat that as logging you bolt on at the end.
So the shape of this project was a bet: design the audit trail and the tenant boundary first, then hang retrieval off them. That ordering is now built. Whether it survives contact with a real deployment is still open, and I do not think code passing its own test suite tells me much about that.
What is actually in the repo.
A list of what the code covers. Which of these has been run against real infrastructure is answered in the roadmap, and that is the file that gets updated.
The decisions behind it.
These are the arguments worth having, and they are the part of this page least likely to be wrong next week.
nomic-embed-text through Ollama, with the BGE and E5 families as an opt-in extra. All of it runs on hardware the customer already owns.
pgvector by default, because most regulated shops already run Postgres and already know how to back it up and audit it. Qdrant sits behind the same interface as an override.
Llama 3, Mistral or Qwen through Ollama. The vLLM adapter is written against the documented API and has never been pointed at a GPU box, so it carries an unverified marker in the code instead of a claim it cannot back.
FastAPI with OpenAI-compatible endpoints, so tooling that already speaks to a cloud provider keeps working after one base URL change. RBAC sits on the same surface.
Filtering is enforced inside the vector store search and delete calls, so there is one choke point rather than a rule every caller has to remember. A new store implementation subclasses the shared contract test. It does not get to write its own.
A knowledge base records the embedding model and dimension it was built with, and a write that does not match is rejected. The alternative is quietly returning neighbours that mean nothing, which is the kind of bug you find six weeks later.
An append-only Postgres table with a per-tenant hash chain. Rows are written after the operation they describe, which the docs say out loud, because it is the sort of detail an auditor finds anyway.
A deterministic faithfulness check as the primary metric, so the evaluation suite itself never has to send your documents to a judge model. Ragas is there for teams who are permitted to use it.
The parts that are meant to make you uncomfortable.
Every one of these is a property of the design. None of them gets fixed by another week of building, so none of them is going to quietly disappear from this page.
- The hash chain makes tampering evident. It does not make it impossible. Anyone with write access to that database can rewrite history and recompute the chain, and there is a test in the repo that does exactly that and still passes verification. A broken chain proves the log was altered. An intact chain proves nothing.
- make demo finishes with two eval cases failing, and they are supposed to. The stand-in embedder is a hashed bag-of-words projection with no semantics in it, so two off-corpus questions score 0.24 and 0.23 while two legitimate ones score 0.22 and 0.25. No score floor separates those. Tuning a threshold until the run went green would be fitting to the fake, so the demo leaves it at zero and prints the reason.
- The deterministic metrics are lexical. They are good at catching a number the source never contained and bad at scoring an answer that says the right thing in different words.
- Retrieval is a single dense-vector pass. There is no hybrid stage, no reranker and no query rewriting.
- Several adapters have only ever met a mocked transport. Which ones those are, and which have been run against a real service, is the second table in the roadmap.
- Running this inside your own perimeter does not satisfy any regulation by itself. It removes one specific data-egress problem and leaves you the rest of the work.
Status lives in one file.
This page carried its own milestone table for about a day, and the table was wrong before the day was over. Five surfaces restating the same status is five chances to publish something that stopped being true, and the repo is the only one of the five I update while I am working.
So the count sits in docs/roadmap.md and everything else points at it. That file carries the milestone list, a section on what the gateway still does not do, and a per-backend table separating what has been run against a real service from what has only been tested against a mock. If it disagrees with this page, believe it.
Who this is being built for.
Banks, NBFCs and insurers in regulated markets. Healthcare organisations handling PHI. Government and defence. Anyone whose security team has already blocked the cloud providers and has been living without retrieval since.
If that describes where you work, open an issue with the constraint you are actually stuck on. I would rather build against a real one than a guessed one, and this is still early enough that the feedback changes the design instead of just annoying me.
Reading it is still more useful than running it.
The argument here is how a retrieval system should be put together when the audit trail and the tenant boundary have to hold up in a review, and that argument is sitting in the repo where you can disagree with it. The next honest step is a real deployment with real documents, and I do not have one of those yet.