“Build or buy?” is usually framed as a software-cost comparison. For AI systems, the more important question is which part of the operating advantage should your organization own?
A company can buy a capable model and still need custom workflow design, data contracts, evaluation, interfaces, and authority controls. It can also build an elegant custom application that recreates a commodity product and becomes expensive to maintain.
The right answer is often a deliberate combination. This framework makes that combination explicit.
1. Define the job before evaluating products
Write the job from trigger to measurable outcome. Include the user, current process, systems involved, decisions made, exceptions, and final operating measure.
A weak definition says, “We need an AI assistant for operations.” A useful definition says, “When a vendor invoice arrives, the system must gather the purchase order and receiving record, identify mismatches, prepare the recommended disposition, route exceptions to the correct owner, and reduce time to a reconciled decision without increasing payment errors.”
The second description can be evaluated. It also reveals that the model is only one part of the product.
2. Separate commodity capability from differentiating logic
List every capability the workflow requires and place it in one of three groups.
Commodity
Capabilities are commodity when many credible products provide them and ownership creates little strategic advantage. Examples can include foundation-model access, generic document parsing, common connectors, authentication, queues, analytics infrastructure, and standard collaboration surfaces.
Configurable
A product may handle the job if its workflow, policy, data mapping, and interface can be configured without fighting the platform. The key test is not whether a demo can be customized. It is whether the configuration model can represent the real exceptions and remain maintainable.
Differentiating
A capability may deserve custom ownership when it encodes proprietary operating knowledge, a unique decision process, data unavailable to competitors, a customer experience central to the brand, or a feedback loop that compounds advantage.
Buy the commodity layer when it meets the contract. Spend custom effort where ownership changes the business outcome.
3. Compare four decisions, not two
A complete review has four possible outcomes.
Buy
Choose a product when it handles the important workflow, meets integration and control requirements, has acceptable operating economics, and does not force the organization to surrender meaningful differentiation.
Build
Choose custom development when the workflow is valuable, specific, and stable enough to define; products cannot satisfy a critical contract; and the organization can own operation after launch.
Combine
Use bought infrastructure or products for commodity capabilities while building the differentiating workflow, policy, evaluation, or experience around them. This is frequently the most practical architecture.
Stop
Do not automate when the baseline is unclear, value is too small, inputs are not observable, decision ownership is missing, or the workflow should be simplified before software is added.
A decision process that cannot recommend “stop” is a procurement ritual, not an evaluation.
4. Evaluate products against the workflow contract
Create a test set before vendor demonstrations. Include normal cases, exceptions, unavailable data, ambiguous inputs, prohibited actions, integration failures, and the performance conditions that matter.
Ask each option to demonstrate the same contract:
- Can it represent the full workflow and its exceptions?
- Which systems can it read and write?
- How are identity and permissions enforced?
- Where does customer data go?
- Which actions require human approval?
- Can behavior be evaluated automatically?
- What is logged and exportable?
- How are prompts, models, and policy versions controlled?
- What happens when the provider, model, or connector fails?
- Can you leave with your data and operating history?
Do not score a product on the best path its sales team prepared. Score the conditions your operators will encounter.
5. Calculate total ownership, not initial implementation
For each option, estimate cost across the expected decision horizon:
- licenses and model usage;
- integration and data work;
- configuration or custom development;
- evaluation and quality assurance;
- security and governance review;
- migration and change management;
- monitoring and incident response;
- vendor-management effort;
- internal product and engineering ownership;
- switching or exit cost.
A custom build is not “one project.” A bought product is not “one subscription.” Both create an operating responsibility.
Model the cost per completed business outcome where possible, not cost per token or seat in isolation.
6. Inspect lock-in at the correct layer
Model portability is only one form of lock-in. The harder dependencies may be:
- proprietary workflow definitions;
- accumulated evaluation data;
- embedded business rules;
- identity and permission mappings;
- agent state and history;
- vendor-specific connectors;
- user habits and approval processes;
- observability available only inside the product.
Record which artifacts can be exported in usable form. A standard model API does not make the whole system portable.
At the same time, do not over-engineer theoretical portability. A replaceable interface is valuable only if the organization has evidence that change is plausible and worth the extra complexity.
7. Use a weighted decision record
Score options against criteria agreed before the review:
- workflow fit;
- expected operating value;
- time to meaningful learning;
- integration fit;
- data and authority control;
- evaluation and observability;
- adoption burden;
- ongoing ownership;
- exit path;
- total cost.
Weight the criteria according to the job. A regulated decision workflow may weight authority and auditability heavily. An internal low-risk assistant may emphasize learning speed and adoption.
Attach evidence and uncertainty to every score. Then write the recommendation, rejected alternatives, assumptions, and events that would trigger a review.
8. Know the signs that custom development is justified
Custom work becomes more defensible when several of these are true:
- the workflow is central to the company’s advantage;
- proprietary data or feedback improves the system over time;
- off-the-shelf products fail a critical exception or control boundary;
- integration is the product rather than a peripheral task;
- the user experience changes customer or operator behavior;
- the organization needs to own evaluation, policy, or deployment;
- value is large enough to fund long-term operation;
- a named team can own the system after launch.
One strong signal is not enough. “Our process is unique” often means it has not been mapped yet.
9. Make the first build a vertical decision slice
If custom development wins, do not begin by recreating the entire workflow. Ship one vertical slice that includes a real trigger, representative data, the key decision, required integration, evaluation, human boundary, and observable outcome.
That slice teaches more than a broad prototype because it crosses the contracts production depends on. It also creates an early opportunity to reverse the build decision if ownership or economics do not hold.
VallySeed’s Custom AI Development work starts with this build-versus-buy decision and keeps commodity capability replaceable where the use case supports it. If the portfolio and priority are still unclear, begin with AI Consulting.