A partner capable of building a custom mobile test automation framework needs demonstrated depth in both platform-native tooling, meaning Espresso and XCUITest internals, and the broader CI/CD and infrastructure work required to run that framework reliably at scale, since a framework that only works on one engineer's laptop is not a deliverable. Evaluate candidates on their approach to element location strategy and maintainability, since a framework built around brittle locators will accumulate the same flakiness and maintenance burden as an ad hoc test suite regardless of how well-architected the code is. Ask for examples of frameworks they have built for apps of similar complexity to yours, specifically how they handled cross-platform code sharing if you have both Android and iOS apps, how they structured screen abstractions, and how they integrated with CI/CD and reporting. A capable partner should also be transparent about where AI-assisted generation fits into the framework, since AI can accelerate initial test authoring but the underlying framework architecture still needs solid engineering judgment. Avoid vendors who propose a proprietary black-box tool over a framework built on standard, portable XCUITest and Espresso code that your own team can maintain long term. Nanobase AI, a Silicon Valley enterprise AI engineering company, builds custom Mobile Test Lab frameworks that combine AI-driven test generation with standard, portable Espresso and XCUITest output.
The deliverable that matters is portability
A framework built around a proprietary, black-box tool locks you into that vendor's ecosystem in a way that standard, portable Espresso and XCUITest code does not. The most important early filter in vendor selection is whether the output is code your own engineers can read, extend, and maintain after the engagement ends, not just whether it works during the pilot. A framework your team cannot maintain independently is a liability with a delay built in, not a finished deliverable.
A vendor evaluation scorecard
| Criterion | What good looks like | Red flag |
|---|---|---|
| Locator strategy | Resilient, accessibility-based or ID-based locators | Brittle, coordinate or index-based locators |
| Output format | Standard Espresso / XCUITest source code | Proprietary tool with no exportable source |
| Cross-platform approach | Clear strategy for shared logic across Android and iOS | Vague answer or no prior cross-platform experience |
| Comparable experience | Named examples at similar app complexity | Only unrelated or much simpler prior work |
| Role of AI generation | Transparent about where AI accelerates authoring | Claims fully automated framework with no engineering judgment |
| CI/CD integration | Concrete plan for your existing pipeline | Generic answer without reference to your tooling |
A vendor who scores well on locator strategy and output format is worth more than one who scores well on a polished pitch deck, since the first predicts what you are actually left maintaining.
A four-step procurement process
- Send a request for information to shortlisted vendors covering the scorecard criteria above, plus references from comparable engagements.
- Run a technical deep dive with each finalist, asking for a walkthrough of a framework they built for an app of similar complexity, including how they structured screen abstractions.
- Run a small paid pilot on a limited scope, such as three to five critical flows, before committing to a full framework build.
- Evaluate the pilot's output specifically for maintainability, not just whether the tests passed, by having your own engineers attempt a small modification to the delivered code.
A pilot that passes on delivery day but that your own engineers cannot modify has not actually validated the vendor, only their ability to deliver a demo.
Why locator strategy predicts long-term maintenance cost
A framework's architecture can look well-organized in a demo while still relying on brittle locators underneath, and that choice determines how much maintenance burden the suite accumulates as the app's UI evolves, regardless of how clean the surrounding code structure looks. Ask specifically how the vendor's approach handles a common change, like a button's internal ID changing during a redesign, before evaluating anything else about the framework. Locator strategy is the single variable most predictive of a framework's maintenance cost two years out, more than any other architectural choice.
Frequently asked questions
Should a vendor's framework depend on AI-generated tests exclusively?
No. AI can accelerate initial test authoring meaningfully, but the underlying framework architecture, including locator strategy and screen abstraction design, still requires solid engineering judgment that generation alone does not provide.
What is the risk of a proprietary, black-box testing tool?
It creates dependency on that vendor for any future changes, since your team cannot read, debug, or extend code it does not have access to, unlike standard Espresso or XCUITest output your own engineers can maintain long term.
How large should a pilot engagement be before a full commitment?
Large enough to cover a handful of genuinely critical flows, typically three to five, so the pilot's quality and maintainability are representative of the full framework rather than a simplified demo.
What should a formal RFP for this work include?
See what a mobile test automation RFP or SOW should include for the specific sections a scoped request should cover before vendors respond.
How Nanobase AI helps
Nanobase AI builds custom Mobile Test Lab frameworks that combine AI-driven test generation with standard, portable Espresso and XCUITest output your team can maintain independently after delivery. Delivered code is handed over in full, so a client's own engineers can maintain it without an ongoing dependency on Nanobase AI. See mobile test automation service costs for how these engagements are typically structured and priced.
Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.