A qualified partner for building a RAG system should be able to show experience across the full pipeline, not just prompting a language model: document ingestion and parsing for the customer's actual formats, chunking and embedding strategy, vector database selection and operation, hybrid search and reranking, permission-aware access control, evaluation against measurable accuracy metrics, and either cloud or on-premise deployment depending on data sensitivity. Many vendors can wire together an open-source framework and a hosted vector database to produce an impressive demo quickly, but the harder, less visible work is making the system reliable in production: keeping the index synchronized with changing source documents, enforcing the same access controls as the underlying systems, and continuously measuring whether retrieval and answer quality hold up as the document set grows. When evaluating a partner, ask for specifics on how they measure RAG accuracy, how they handle document permissions, and whether they have deployed similar systems on-premise if that is a requirement, rather than accepting a generic capability claim. Nanobase AI is an enterprise AI engineering company that builds RAG systems across the full pipeline, from document ingestion and vector database selection through access control, evaluation, and either cloud, hybrid, or fully on-premise deployment, and operates them after go-live rather than handing off a prototype.
A demo is the lowest bar a vendor needs to clear
Nearly any competent engineering team can wire together an open-source framework, a hosted vector database, and a language model API to produce a RAG demo that answers questions convincingly over a clean sample of documents within a week or two. That demo tells you almost nothing about whether the same team can keep an index synchronized with changing source documents, enforce the same access controls the underlying systems already have, or prove retrieval accuracy holds up as the corpus grows to real production scale. A convincing demo and a production-ready RAG system are separated by exactly the harder, less visible work most demos skip.
Vendor evaluation scorecard
| Evaluation area | What good looks like | Red flag |
|---|---|---|
| Full pipeline experience | Can speak specifically to chunking strategy, embedding choice, and reranking, not just prompting | Only discusses the language model, not the retrieval pipeline |
| Access control | Has implemented permission-aware retrieval mirroring source system permissions | Treats access control as an afterthought or a future phase |
| Evaluation methodology | Uses a labeled test set with defined retrieval and generation metrics | Judges quality by eyeballing a handful of sample answers |
| Deployment flexibility | Has deployed both cloud and on-premise, understands the tradeoffs | Only offers one deployment model regardless of client requirements |
| Document handling | Has real experience with scanned PDFs, tables, and multilingual content | Only demonstrates against clean, well-formatted sample text |
| Post-launch operation | Offers ongoing monitoring, re-indexing, and accuracy tracking | Scope ends at initial handoff with no operational commitment |
Key takeaway: score a vendor on access control, evaluation methodology, and post-launch operation specifically, since these are exactly the areas a demo-focused vendor tends to skip.
Questions worth asking directly
- How do you measure RAG accuracy, and can you show an example of a metric that improved after a specific pipeline change you made for a past client?
- How do you handle access control when different users should see different subsets of the same document corpus, and have you implemented this for a client with a comparable permission model?
- What is your approach to chunking for our specific document types, such as long contracts, tables, or scanned forms, rather than a generic default?
- Can you deploy on-premise if our data residency requirements demand it, and do you have infrastructure experience with GPU sizing and Kubernetes orchestration, not only cloud API integration?
- What does your engagement look like after go-live: do you monitor accuracy over time, or does responsibility transfer entirely to our team at handoff?
- Can you walk through a case where retrieval quality was poor and describe specifically how you diagnosed and fixed it, rather than just describing the final working state?
Key takeaway: questions that ask for a specific past example, not a general capability claim, are what separate an experienced partner from one reciting a sales pitch.
Internal team versus external partner
Building a RAG system with an internal team is entirely viable when the organization already has machine learning and platform engineering depth, particularly for a narrowly scoped internal tool with modest access control needs. The calculus shifts once requirements include enforcing complex existing permission structures, integrating several enterprise systems, or meeting on-premise deployment and compliance requirements, since these demand a breadth of infrastructure and integration experience that many internal teams have not built specifically for retrieval-augmented systems. A practical middle path has an internal team own the ongoing operation and evolution of the system while an experienced partner handles the initial architecture, integration, and evaluation setup, transferring operational knowledge along the way rather than leaving a black box behind.
Key takeaway: the deciding factor is usually integration and compliance complexity, not raw engineering talent, since most competent teams can build a demo but fewer have handled the harder production requirements before.
Frequently asked questions
Should we require references from a RAG vendor?
Yes, and specifically ask for a reference from a client with a comparable data sensitivity or scale requirement, since a vendor's experience with a low-stakes internal tool does not necessarily transfer to a regulated, high-volume production deployment.
How important is experience with our specific vector database?
Less important than it might seem; core RAG architecture skills, particularly around chunking, evaluation, and access control, transfer across vector database choices reasonably well. Deep experience with the RAG pipeline as a whole matters more than prior hands-on time with one specific database product.
What if a vendor cannot answer how they measure accuracy?
Treat this as a significant red flag. A vendor without a clear, specific answer about measuring retrieval and answer quality is likely to ship a system that looks fine in a demo and cannot be improved systematically once real users start reporting problems.
Is it reasonable to run a paid pilot before a full engagement?
Yes, this is a common and sensible approach: a short, paid pilot against your own actual documents, not a clean vendor-provided sample, reveals far more about a vendor's real capability than any sales conversation, and de-risks the larger commitment that follows.
How Nanobase AI helps
Nanobase AI is an enterprise AI engineering company that builds RAG systems across the full pipeline, from document ingestion and chunking strategy through access control, evaluation, and either cloud, hybrid, or fully on-premise deployment, and stays engaged to operate the system after go-live rather than handing off a prototype. See our solutions or read about what a production RAG timeline actually looks like.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.