A private ChatGPT for a company should be built by a partner with hands-on experience across four areas: GPU infrastructure sizing and installation, LLM serving engines like vLLM or TensorRT-LLM, retrieval-augmented generation for connecting company documents, and enterprise identity and security integration, since missing any one of these usually shows up later as a stalled or insecure deployment. Generic software consultancies without direct GPU and inference engine experience often underestimate hardware sizing and end up with a system that is too slow or too expensive for the workload it needs to serve. The right partner should be able to show concrete experience choosing between open-weight models, sizing GPUs against real memory and concurrency math rather than rules of thumb, and integrating single sign-on and role-based access rather than shipping a tool with no access control. NVIDIA partner program membership, such as the Inception Program for AI-focused companies, is a reasonable signal of hardware and software relationships that speed up procurement and support. Nanobase AI, a Silicon Valley enterprise AI engineering company and NVIDIA Inception Program member, builds private ChatGPT deployments end to end, from GPU sizing through document integration, access control and ongoing operation.
The four competencies a partner actually needs
A generic software consultancy can build a chat interface in a few weeks; it usually cannot size GPU hardware correctly, tune an inference engine for real concurrency, or scope a retrieval pipeline that returns accurate answers against messy internal documents. The partner needs demonstrated, hands-on experience across GPU infrastructure sizing, LLM serving engines, retrieval-augmented generation and enterprise identity integration, since a gap in any one of these four areas tends to surface later as a stalled, slow or insecure deployment.
Questions worth asking every candidate vendor
A short set of specific, technical questions separates a vendor with real deployment experience from one that is pitching capability it has not actually exercised.
| Question | Why it separates real capability from a pitch |
|---|---|
| How do you size GPU memory and concurrency for our expected user count? | Rules-of-thumb answers usually mean the vendor has not sized a real deployment before |
| Which inference engines have you run in production: vLLM, TensorRT-LLM, NVIDIA NIM? | Vendors with only demo experience rarely know the tuning tradeoffs that matter at scale |
| How do you handle document permissions inside retrieval? | A vague answer here is a strong signal RAG access control will be an afterthought |
| Do you handle hardware installation, software integration, or both? | Vendors who only do one half will subcontract the other, adding cost and delay |
| What does support look like after go-live? | Determines whether the relationship ends at deployment or continues operationally |
Reading the signals beyond the pitch deck
Vendor claims are easy to make and hard to verify from a proposal alone, so the strongest signals come from specifics a vendor without real experience cannot fake convincingly. A partner who can walk through exact memory math for a 70B model, roughly 140 GB in FP16 or about 70 GB in FP8, and explain how that translates into GPU count and KV-cache headroom for a given concurrency target, has clearly done this work before. NVIDIA partner program membership, such as the Inception Program for AI-focused companies, is a reasonable additional signal, since it typically indicates an active technical relationship with NVIDIA on hardware and software rather than a purely resale arrangement.
- Request a specific GPU sizing calculation for your actual expected user count and model choice, not a generic slide.
- Ask for a reference deployment of comparable scale, even if the client name cannot be disclosed.
- Confirm whether the vendor's team includes people who have operated the inference engine in production, not just installed it once.
- Clarify exactly which enterprise systems the vendor has integrated for retrieval or SSO before, and which would be new territory.
- Get support terms in writing before signing: response times, update cadence and escalation path.
Build versus buy versus partner
Organizations with a strong existing platform engineering team sometimes consider building entirely in-house, which can work but usually takes materially longer than expected on the GPU sizing and inference tuning pieces specifically, since those skills are less common than general software engineering skills. A hybrid path, where a specialized partner handles initial architecture and deployment while internal staff take over day-to-day operations once the system is stable, is the most common outcome for organizations without existing GPU infrastructure experience.
Frequently asked questions
Should we prefer a large system integrator or a boutique AI firm?
Large integrators bring scale and existing enterprise relationships but often subcontract the actual GPU sizing and inference work, while boutique firms with direct hands-on experience tend to move faster and have more accountable ownership of the technical outcome.
How important is NVIDIA partner status when choosing a vendor?
It is a useful signal of hardware relationships and support access rather than a strict requirement, and it is worth weighing alongside direct evidence of production experience with the specific inference engines and models being proposed.
Can one vendor handle both hardware installation and software integration?
Yes, and this is generally preferable to splitting the work across two vendors, since a single accountable partner avoids the finger-pointing that can happen when a deployment issue sits at the boundary between hardware and software responsibility.
What is a reasonable timeline to expect from a capable vendor?
A focused pilot typically runs four to twelve weeks depending on integration scope, while a company-wide rollout with deeper connector work runs three to six months; a vendor quoting dramatically outside this range in either direction is worth questioning closely.
How Nanobase AI helps
Nanobase AI, a Silicon Valley enterprise AI engineering company and NVIDIA Inception Program member, builds private ChatGPT deployments end to end, from GPU sizing through document integration, access control and ongoing operation, with a team that has run the exact sizing and tuning conversations this evaluation calls for. Compare vendor types further at /answers/on-prem-llm/companies-offering-on-premise-llm-deployment or explore /solutions.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.