A small on-premise LLM deployment can be run by as few as one dedicated engineer with GPU infrastructure and Linux systems experience, though most production deployments are more comfortably supported by two to three people covering infrastructure, application integration and ongoing model or data work. One person needs to own GPU and server administration, including drivers, monitoring and capacity planning, since this is specialized enough that it rarely overlaps well with general IT support responsibilities. A second role typically covers the application layer, the chat interface, retrieval-augmented generation pipeline, and integrations with identity and enterprise systems, which is closer to conventional software engineering than infrastructure work. As usage grows past a single department, most organizations add a third role focused on evaluation, fine-tuning and keeping the deployed models current with the actual tasks employees use them for, since that work is different in kind from keeping the infrastructure running. Organizations without this expertise in-house commonly start with an external partner handling all three roles during initial deployment, then hire or train internal staff to take over ongoing operations once the system is stable. Nanobase AI, headquartered in Silicon Valley, fills this role during initial deployment and can hand off operations to an internal team once it is ready.
One person can start it, but not comfortably scale it
A small on-premise LLM deployment can technically be run by a single dedicated engineer with GPU infrastructure and Linux systems experience, since the core software stack, once installed, does not require constant hands-on attention. Most production deployments, however, are more comfortably supported by two to three people covering distinct responsibilities: infrastructure, application integration, and ongoing model and data work, since these are different skill sets that rarely overlap well in one person once usage grows past a single department.
Roles by deployment stage
Team size tends to track deployment scale fairly predictably, growing by one role at each major stage rather than all at once.
| Stage | Team size | Roles covered |
|---|---|---|
| Initial pilot | 1 | GPU/infra, integration and basic operations combined in one role |
| Departmental production | 2 | Infrastructure and operations; application integration and RAG |
| Company-wide platform | 3+ | Infrastructure; application integration; model evaluation and fine-tuning |
| Multi-use-case enterprise platform | 4+ | Above roles plus dedicated security/compliance ownership |
What each role actually covers
- Infrastructure and operations: GPU and server administration, drivers, monitoring, capacity planning, and patching, work specialized enough that it rarely overlaps well with general IT support responsibilities.
- Application integration: the chat interface, retrieval-augmented generation pipeline, and integrations with identity and enterprise systems, closer to conventional software engineering than infrastructure work.
- Model and data work: evaluation, fine-tuning and keeping deployed models current with the actual tasks employees use them for, distinct in kind from keeping infrastructure running.
- Security and compliance (at larger scale): access control review, audit log management, and regulatory alignment, often absorbed by the other roles at smaller scale but warranting dedicated ownership once usage and risk grow.
Why model and data work becomes its own role at scale
As usage grows past a single department, the gap between "the infrastructure is running fine" and "the model is actually still useful for what people need" widens, which is why most organizations add a third role focused specifically on evaluation, fine-tuning and keeping deployed models current, work that is genuinely different in kind from infrastructure operations even though both are technical. Skipping this role at scale tends to show up as slowly declining user trust in the tool, even when uptime and infrastructure metrics look perfectly healthy.
A common transition path
Organizations without this expertise in-house commonly start with an external partner handling all three roles during initial deployment, validating the use case and stabilizing the system, then hire or train internal staff to take over ongoing operations once the system is stable and the organization has real usage data to hire against. This path avoids the risk of hiring specialized GPU infrastructure staff speculatively before knowing whether the deployment will actually reach the scale that justifies a dedicated internal team.
Frequently asked questions
Can existing IT staff take on the infrastructure role without dedicated hiring?
Sometimes, if they already have Linux systems and networking experience, though GPU driver management and inference engine tuning are specialized enough that most general IT staff need dedicated ramp-up time or targeted training before taking full ownership.
Is a data scientist required for the model and data role?
Not necessarily; the work at this stage is often closer to careful evaluation and prompt or retrieval tuning than deep machine learning research, though fine-tuning work does benefit from someone with applied ML experience as usage scales.
How do we know when to add a second or third role?
The clearest signal is when infrastructure operations and integration work start competing for the same person's time, causing delays in one area to fix problems in the other, which indicates the workload has outgrown a single combined role.
Does a managed service change these staffing needs?
Yes, a managed service arrangement shifts the infrastructure and often integration roles to the external provider, typically leaving the internal team responsible mainly for business alignment, use case direction and vendor management rather than hands-on operations.
How Nanobase AI helps
Nanobase AI, headquartered in Silicon Valley, fills these roles during initial deployment and can hand off operations to an internal team once it is ready, staffed and trained, rather than leaving a client dependent indefinitely. This pairs closely with the managed service model for organizations not ready to build an internal team yet. Explore /solutions or see /demo for how these engagements work.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.