Technical methodology / model v1.0.0

Build vs Buy AI Agent Scorecard Methodology

This document defines the exact scoring, evidence, constraint and three-year TCO behavior used by the public tool.

Published
Last source verification
Model version
1.0.0
Export schema
1.0.0

Purpose and scope

The scorecard records a sourcing posture for one defined AI agent use case. It does not compare named vendors, estimate project success or certify an architecture. Teams should run it after they have a workflow, data boundary, integration inventory, outcome owner and initial operating assumptions.

The model has three outputs: a direction, an evidence status and a constraint status. Keeping them separate prevents a strong-looking number from hiding missing facts or a mandatory blocker.

Interpretation boundary: the decision index is an authored decision-support scale. It is not a probability, confidence score, market benchmark, forecast or financial model.

Scoring contract

Twelve criteria, four answer states and 31 signed points

Answer values

Buy
-1 multiplied by the criterion weight
Hybrid
0 multiplied by the criterion weight
Build
+1 multiplied by the criterion weight
Unknown
null; the criterion is removed from known weight

Weights

Seven criteria carry weight 3: strategic differentiation, workflow uniqueness, data control, integration depth, time to validated value, team capability and lifecycle ownership. Five carry weight 2: capital shape, scale economics, portability, evaluation control and compliance path. The absolute total is 31.

index = round-half-away-from-zero(100 x sum(answer value x weight) / known weight)

Normalization uses known weight, not total weight. Unknown answers therefore reduce evidence coverage without pulling the direction toward hybrid. If nothing is known, the index is displayed as zero but the outcome remains unresolved.

Outcome thresholds

Version 1.0.0 decision thresholds
IndexProvisional direction
-100 through -30BUY
-29 through +29HYBRID
+30 through +100BUILD

The thresholds make the editorial model reproducible. They have not been calibrated against a dataset of completed projects and should not be presented as universal success boundaries.

Decision-boundary warning

A result within seven index points of -30 or +30 receives a warning. This is a sensitivity prompt, not a confidence interval. Its practical meaning is that one high-weight answer may change the sourcing direction.

The export records the distance to the nearest threshold so reviewers can identify which evidence deserves challenge before approval.

Uncertainty contract

Evidence readiness can withhold a numeric answer

The model returns INSUFFICIENT_DECISION_EVIDENCE when known-weight coverage is below 75% or when any of four gate-critical criteria is unknown. Those critical criteria are data control, team capability, lifecycle ownership and compliance path.

Coverage is calculated as known weight divided by total weight. The record still shows the direction among known answers as a provisional outcome, but the public outcome is UNRESOLVED. This preserves useful information without presenting partial evidence as a decision.

Why these four gates are required

  • Data control can disqualify a vendor path even when every feature fits.
  • Team capability can disqualify a custom build even when it would be strategically attractive.
  • Lifecycle ownership determines whether evaluation, releases and incidents can be operated.
  • Compliance path determines whether the required approvals and records can be produced.

Constraint override

Two build-leaning answers can block a pure buy outcome: no approved vendor route meets the mandatory data-control requirement, or no current vendor route meets a mandatory compliance gate. Two buy-leaning answers can block a pure build outcome: no durable internal build team, or no internal lifecycle capacity.

When one of those selected answers conflicts with the provisional direction, the record returns CONSTRAINT_REVIEW_REQUIRED and uses HYBRID as the working boundary. That does not prove hybrid is feasible. It tells the team which ownership boundary must be redesigned or which evidence must change.

Cost contract

Three-year TCO separates cadence before addition

three-year TCO = one-time total + (annual total x 3) + exit reserve

Cost categories and cadence in model v1.0.0
CategoryCadenceCoverage
Discovery and architectureOne-timeWorkflow, data, risk, product and architecture discovery
Implementation and integrationOne-timeConfiguration or development, environments, testing and release work
Data preparation and migrationOne-timeCleanup, permissions, migration, indexing and initial evaluation data
Licenses, models and infrastructureAnnualPlatform, model usage, cloud, storage, networking and observability services
Internal product and operations teamAnnualProduct, engineering, data, platform and incident ownership retained internally
Evaluation, security and complianceAnnualTest maintenance, red teaming, reviews, controls and audit evidence
Change, training and supportAnnualUser enablement, support, workflow change and system updates
Exit and switching reserveEnd of horizonExport, migration, termination help, retraining and parallel run

The bundled values are an editable worked example. They are not company pricing, vendor quotes or market averages. The calculator rejects negative and non-finite values, preserves each category in exports and breaks ties in the stable order build, buy, hybrid.

Export contract

JSON, CSV and Markdown exports include model version, evidence status, outcome, provisional direction, index, signed points, known weight, total weight, coverage, unknown inputs, blocking constraints, criterion answers and TCO totals.

The JSON record also snapshots claims and source metadata. A later update to a public registry does not silently change the decision record already exported.

State and privacy

The site computes in the browser. It does not send answers to a server or preserve them in local storage, session storage, cookies or the URL. Any edit marks the previous result stale and disables exports until recalculation.

Evidence ledger

Sources and their limits

  1. NIST AI Risk Management FrameworkSupports full-lifecycle risk and acquisition questions. It is voluntary and does not define sourcing thresholds.
  2. NIST AI 600-1 Generative AI ProfileSupports generative AI lifecycle risk practices. It does not validate weights or economics.
  3. FinOps Foundation terminologyDefines broad TCO and unit economics. It supplies no agent price ranges.
  4. KPMG Agentic AI UntangledSupports readiness dimensions and has a disclosed consultancy interest.
  5. The Buy-or-Build Decision, RevisitedA conceptual preprint that supports factor coverage, not empirical calibration.

Machine-readable limits, verification dates and claim bindings are available in sources.json and claims.json.

Authorship, review and change governance

Publisher
Pharos Production Editorial Team
Independent review
Not performed for release 1.0.0
Drafting
AI-assisted drafting with source and deterministic QA
Correction policy
Material corrections are dated in the changelog; behavior changes increment the model version.

Recheck the source registry and assumptions at least quarterly. Reissue the model sooner when a source is superseded, a regulation changes, the scoring behavior changes or release QA discovers a material defect.

Return to the scorecard