AI Agent Build vs. Buy: What Should You Own?
AI agent Build vs. Buy is often reduced to “Should we train a model or call an API?” But the model is only one layer. Workflow rules, data, permissions, evaluation sets, user experience, and the ability to change vendors may matter for much longer.
This is not a binary decision. Own the workflow logic that creates differentiation, buy commodity capability, and keep control of data, permissions, evaluation, and exit options in every architecture.
Distinguish Four Options
| Option | Description | Speed | Control and Flexibility | Best Fit |
|---|---|---|---|---|
| Packaged SaaS | Vendor provides UI, workflow, model, and operation | Fastest | Low to medium | Standard support and productivity work |
| Low-code or managed agent | Configure work and integrations on a vendor platform | Fast | Medium | Rapid integration with some customization |
| Managed model/API plus custom app | Buy model and infrastructure; own UX, workflow, and evals | Medium | Medium to high | Workflow logic and data differentiate |
| Self-operated stack | Operate model, serving, and orchestration | Slow | Highest | Scale, regulation, or core IP justifies the cost |
Most companies use a portfolio rather than one answer. Buy the commodity and build the differentiation.
A 100-Point Decision Matrix
Score each criterion from one, a Buy signal, to five, a Build signal. These weights are a starting point for discussion.
| Criterion | Weight | 1: Buy Signal | 3: Hybrid | 5: Build Signal |
|---|---|---|---|---|
| Strategic differentiation | 20 | Generic back office | Some experience differentiation | Core product or operating IP |
| Workflow specificity | 15 | Industry-standard process | Some approvals and exceptions | Complex proprietary rules and state |
| Integration depth | 10 | Standard connector | Several internal APIs | Legacy, real-time, multiple systems |
| Data and regulatory control | 15 | General data and standard terms acceptable | Regional and retention constraints | Highly sensitive, sovereignty, audit |
| Evaluation and operating capability | 10 | No dedicated team | Shared with a partner | Internal eval, SRE, and security |
| Time to market | 10 | Needed immediately | Needed this quarter | Long-term investment possible |
| Three-year TCO and scale | 10 | Small and variable demand | Breakeven unclear | Large stable demand and unit advantage |
| Portability and lock-in | 10 | Lock-in acceptable | Data portability required | Vendor dependence is strategic risk |
Interpretation
- 0–39: Buy first. Strengthen vendor diligence and the data exit plan.
- 40–69: Prefer Hybrid. Own workflow logic, data, and evaluation; buy models and infrastructure.
- 70–100: Consider Build, but only with real staffing, evaluation, and 24/7 operating budget.
These ranges are prompts for discussion, not a standard. Regulatory prohibitions or mandatory data-location constraints remain separate gates.
Put Hidden Costs in the Same Table
| Cost or Responsibility | Easy to Miss in Buy | Easy to Miss in Build |
|---|---|---|
| Initial | Integration, migration, security review | Data, eval set, and platform construction |
| Recurring | Seats, usage, outcome fees, price changes | Inference, storage, observability, on-call, retention |
| Quality | Vendor-wide metrics versus your task | Regression evals, model replacement, drift |
| Risk | Subprocessors, transfer, changing terms | Patching, vulnerabilities, incident response |
| Exit | Export of data, logs, and prompts | Technical debt and key-person dependency |
Three-year TCO includes more than licenses and engineering salaries.
TCO = license and inference + build and integration + data preparation + evaluation and monitoring + security and compliance + human review + failure and recovery + change and training + exit and migration
Twelve Questions Before Buying
- Is customer data used for training, and what is the default?
- Where is data stored, for how long, and under what deletion and transfer rules?
- How are changes to subprocessors and model providers disclosed?
- Are roles, least privilege, SSO, SCIM, and audit logs supported?
- Can approvals and amount or target limits be applied before tool execution?
- Can customer eval sets run regression tests across versions?
- What are the SLA, support hours, and incident-notification deadline?
- Can the vendor unilaterally change models, price, or functionality?
- At what level can usage, cost, success, and human intervention be exported?
- Can data, memory, logs, and prompts be exported in standard formats?
- Is there evidence of deletion and a continuity plan after termination?
- Do relevant certifications cover the service you will actually use?
The UK government’s Guidelines for AI procurement recognizes that AI may be built from scratch, bought off the shelf, or added to existing systems. It emphasizes data assessment before procurement, multidisciplinary decisions, and ongoing contract management.
Minimum Conditions for Build
- A product or workflow owner and at least a 12-month budget
- Ability to create a real-work eval set and graders
- Operators for security, permissions, observability, and incident response
- Abstractions that preserve the workflow contract and evals when models change
- Quantified differentiation or long-term unit economics that justify construction
Anthropic’s Building Effective Agents recommends beginning with the simplest solution and adding complexity only when needed. Build does not automatically mean multiple agents or a self-trained model.
Three Illustrative Decisions
- Internal meeting summaries: Buy packaged SaaS; scrutinize data, retention, and access.
- Industry-specific customer support: Hybrid; use managed models while owning knowledge, evals, and approvals.
- A product’s core automation engine: Build the workflow orchestration while keeping models replaceable across providers.
The conclusion is neither “always build” nor “always buy proven SaaS.” Own the workflow logic, data, evaluation, and exit—not necessarily the model.