Skip to main content
Executive Guide | Part 2

Executive Guide to Enterprise AI Implementation

A Design, Build and Scale operating guide for moving enterprise AI from a bounded opportunity to an accountable production service.

Accountable delivery map

Evidence and ownership accumulate through every stage.

  1. 01

    Design

    Define the service, authority boundary and evidence required.

  2. 02

    Build

    Prove behaviour, integration, control and recovery under realistic conditions.

  3. 03

    Scale

    Operate the service and govern change across the portfolio.

Persistent decision lenses

  • Value
  • People
  • Evidence
  • Control
Executive answer

Enterprise AI becomes real when the service is accountable

A production system needs a named outcome, process boundary and owner. Capability selection follows consequence, authority and evidence requirements. Design decisions become acceptance criteria and operating controls, while scale is earned through evidence rather than inferred from a successful demonstration.

Development edition August 2025 · First 36-page series edition August 2026

What changes from an AI pilot to a production service
Decision areaPilotAccountable production service
OwnershipProject sponsorNamed service owner and accountable executive
EvidenceDemonstration resultsVersioned evaluation against representative and adverse scenarios
Data and toolsSelected test accessGoverned identities, permissions, lineage and action boundaries
FailureManual recoveryDesigned refusal, hand-off, rollback and incident response
ChangeInformal iterationMaterial-change triggers, re-evaluation and release records

Part 1 prerequisite

Start with the capability boundary

Part 1 defines the Humint Labs 4-Pillar AI Maturity Framework. Part 2 begins when the organisation has framed an outcome and needs to implement the appropriate capability as an accountable service.

Review the 4-Pillar framework
01 | Design

Design the service, authority and evidence

Design is the operating foundation. It defines what the system is allowed to change, how people remain informed and in control, and what must be proved before release.

Design the opportunity and operating boundary

Begin with a measurable service outcome, the affected people and process, the current baseline, the consequence of error and the executive who owns the decision.

The opportunity brief should state what changes, what remains outside scope and what would make the investment unjustified. Model or platform selection follows this decision, not the other way around.

Decision evidence: Opportunity brief, baseline, process boundary, decision criteria and owner map.

Design the service before the system

Map the customer or employee journey, frontstage experience, backstage work, policy decisions, systems of record and measurement points as one service.

A service blueprint exposes work that automation can otherwise displace into complaints, exceptions, manual review or another channel. It also makes integration and operating ownership visible before implementation.

Decision evidence: Service blueprint, journey measures, system boundaries and exception ownership.

Design behaviour people can understand and control

Specify disclosure, permission, confidence, refusal, repair, escalation and human override for every consequential interaction or action.

Agentic Experience Design turns autonomy into an explicit experience and operating decision. People and operators need to understand status, intervene safely and retain the right context at hand-off.

Decision evidence: Interaction flows, repair paths, authority matrix, operator view and hand-off contract.

Define evidence before the system can impress you

Agree representative and adverse scenarios, outcome measures, quality thresholds, intervention expectations, accessibility and recovery criteria before building the prototype.

This prevents a persuasive demonstration from becoming the acceptance test. It also establishes a reproducible basis for choosing the least autonomous capability that can deliver the outcome.

Decision evidence: Evaluation plan, scenario set, thresholds, risk appetite and stop conditions.

Set the authority boundary

Classify the system's role as advise, recommend, prepare, act with approval or act within delegated bounds. Consequence, reversibility and evidence determine the permitted authority.

  1. 01

    Advise

  2. 02

    Recommend

  3. 03

    Prepare

  4. 04

    Act with approval

  5. 05

    Act within bounds

Specialist depth: Design services and Agentic Experience Design.

02 | Build

Build a system that can be proved and changed safely

Implementation connects service intent to architecture, evaluation, control, recovery and a declared production decision.

Build evidence, not a demonstration

Exercise the complete service against expected work, adverse conditions, policy boundaries, tool calls, retrieval, hand-offs and recovery using a recorded configuration.

The evaluation harness must measure the business outcome as well as quality, control failures, latency, cost and human intervention. Results should be reproducible and comparable across model or system changes.

Build an architecture that can change safely

Treat the model as one replaceable component within routing, retrieval, workflow, identity, policy, observability, state and recovery layers.

Record provider and model dependencies, fallback paths, knowledge boundaries, retention, system-of-record writes and the conditions under which the service degrades, stops or returns control to a person.

Keep consequential authority outside the model

Policy checks, permissions and protected-action approvals should be deterministic controls around the model rather than instructions the model can reinterpret.

Separate user identity, system identity, data authority, tool authority and approval authority. Apply least privilege, confirmation, idempotency, rate limits and trusted return-channel checks to consequential actions.

Release against declared gates

A production decision requires an approved scope and owner, architecture and dependency record, evaluation results, control acceptance, recovery proof, support model and rollback decision.

Known limitations belong in the release record. Privacy, security, accessibility, operations and human hand-off are release evidence, not work deferred until after launch.

Enterprise architecture

Architecture is the operating boundary made executable

Enterprise AI is not a model connected to a user interface. It is a service architecture that carries identity, policy, data authority, state, recovery and evidence through every action.

Accountable service architecture

Control and evidence cross every layer. No single platform owns the complete operating decision.

  1. 01Experience and channels
  2. 02Identity, policy and authority
  3. 03Orchestration and state
  4. 04Models, retrieval and tools
  5. 05Data and systems of record
  6. 06Observability and operations
Human controlEvidenceRecovery
Layer 01

Experience and channels

Where people interact, understand system status, give permission, correct work and receive a remedy.

Evidence: Interaction flows, accessibility review, disclosure, hand-off and repair tests.

Layer 02

Identity, policy and authority

Which person or system may access data, invoke a tool, approve an action or act within delegated bounds.

Evidence: Identity mapping, permission tests, policy decisions, protected-action logs and approval records.

Layer 03

Orchestration and state

How workflows, agents, queues, memory and human tasks coordinate, pause, retry and recover.

Evidence: State model, idempotency tests, exception ownership, recovery and compensating-action proof.

Layer 04

Models, retrieval and tools

Which replaceable components generate, retrieve, reason or act, and how routing and fallback operate.

Evidence: Model and tool inventory, evaluation results, routing record, provenance and fallback tests.

Layer 05

Data and systems of record

What data can be read, written, retained or derived, with lineage, classification and purpose intact.

Evidence: Data flow, source authority, retention, access inheritance, write controls and reconciliation.

Layer 06

Observability and operations

How outcomes, quality, control, cost, customer impact, incidents and material change remain visible.

Evidence: Telemetry map, scorecard, alerts, runbooks, incident exercises and lifecycle review records.

Evaluation evidence

Evaluate the service, not only the model

A benchmark score cannot prove that a service will deliver the intended outcome, respect authority or recover safely. Evaluation must connect model behaviour to the real process, people, controls and operating conditions.

Evaluation dimensions, required evidence and accountable owners
Evaluation dimensionWhat must be measuredAccountable evidence owner
Business and service outcomeCompletion, time, rework, customer effort and outcome against the declared baseline.Business or service owner
Task and model qualityCorrectness, groundedness, completeness, calibration and consistency for the actual task.AI product and evaluation lead
Authority and tool usePermission enforcement, valid tool selection, protected actions, confirmation and provenance.Service owner, security and control owner
Failure and recoveryRefusal, timeout, partial failure, bad retrieval, dependency loss, escalation, rollback and repair.Engineering and operations
People and experienceComprehension, accessibility, disclosure, intervention, hand-off quality, complaints and remedy.Design, accessibility and service operations
Operational viabilityLatency, throughput, cost per completed unit, support effort, capacity and provider concentration.Platform, finance and service owner

Representative scenarios

Real distributions, channels, languages, process states and user needs from the intended service.

Adverse scenarios

Prompt injection, bad retrieval, ambiguous authority, dependency loss, unsafe requests and compound failure.

Recorded configuration

Model, prompts, policies, retrieval corpus, tools, data versions, environment and evaluation workload.

90-day decision path

Use time boxes to force decisions, not promise production

Ninety days is a disciplined path to a release, redesign or stop decision for a bounded service slice. It is not a universal promise that every enterprise AI service can reach production in one quarter.

  1. Days 1-3001

    Design the decision

    Confirm the outcome, baseline, affected people, service boundary, owner, authority class and evidence plan. Map the current journey and systems.

    Decision gate: Do not build until the organisation can state what success, harm, intervention and stop conditions mean.

  2. Days 31-6002

    Build and challenge the system

    Implement the narrowest end-to-end service slice. Exercise representative, edge and adverse scenarios across data, tools, hand-offs and recovery.

    Decision gate: Do not expand scope while critical scenarios, protected actions or failure ownership remain unresolved.

  3. Days 61-9003

    Release or stop with evidence

    Complete operational acceptance, runbooks, monitoring, training, support, rollback and executive review. Record limitations and the next decision.

    Decision gate: Release only when accountable owners accept the evidence and retained risk. A stop or redesign decision is a valid outcome.

03 | Scale

Scale accountability with the service

Scale means governing a changing portfolio while outcomes, controls, customer impact and exit readiness remain visible.

Scale the service, not just the workload

Assign service, platform, model, data, risk, support and incident responsibilities before demand grows or authority expands.

The operating model should distinguish day-to-day product ownership from executive accountability and independent control. A universal organisation chart is less useful than explicit decision rights.

Govern the portfolio as dependencies change

Maintain an inventory of workloads, owners, models, providers, data, tools, authority levels, replacement options, lifecycle signals and concentration exposure.

Measure cost per completed unit and route distribution, not only token or licence cost. Exit readiness is an operating capability, not a procurement clause filed away after selection.

Monitor outcomes, controls and customer impact

Connect business outcomes with system quality and control evidence so a technically healthy system cannot conceal poor service or customer harm.

Complaints, overrides, near misses, incidents and recovery outcomes should update design artefacts, evaluation scenarios and release gates. Monitoring is evidence for the next decision.

Treat material change as a new decision

Reassess changes to models, prompts, tools, data, authority, users, jurisdictions, vendors or operating conditions when they can alter outcomes or controls.

Version the evidence, preserve prior decisions and define substitution, rollback and decommissioning paths. A model update is not automatically a routine software change.

What accountable scale looks like

Use an evidence scorecard rather than one synthetic maturity score that can conceal a failed control or poor customer outcome.

Value

Outcome against baseline; cost per completed unit; benefits owner

Service

Completion, rework, hand-off, complaints and recovery

Experience

Comprehension, accessibility, disclosure, control and remedy

System

Quality, reliability, latency, capacity and provider dependency

Control

Policy breaches, interventions, incidents and assurance findings

Lifecycle

Change exposure, replacement readiness and evidence currency

Operating model

Put decision rights beside the evidence

A centre of excellence can set standards and supply shared capability, but it cannot own every service outcome. Accountability must remain with the leaders who control the process, customer impact and operating risk.

Enterprise AI roles, decision rights and evidence records
RoleDecision rightEvidence owned
Accountable executiveOutcome, risk appetite and material expansionBusiness case, decision record and management review
Service ownerEnd-to-end service performance, customer impact and operational acceptanceScorecard, complaints, incidents, backlog and release recommendation
AI product and designCapability boundary, experience, evaluation intent and improvementOpportunity brief, service blueprint, interaction flows and evaluation plan
Platform and engineeringArchitecture, delivery, reliability, security engineering and recoveryArchitecture record, dependency inventory, test evidence, runbook and rollback
Data and model stewardshipData authority, lineage, quality, model selection and lifecycle signalsData flow, model card, source register, routing and substitution evidence
Risk, legal, privacy and securityIndependent challenge, obligation mapping and control acceptanceImpact assessment, control tests, issues, conditions and assurance record
Portfolio decisions

Prioritise evidence-rich opportunities, not impressive demos

Business value and technical feasibility are necessary but incomplete. The portfolio also needs service ownership, evidence, reversibility, customer impact and platform leverage.

01

Outcome value

Material service, revenue, cost, risk or workforce outcome with a credible baseline.

02

Service fit

A bounded process with clear users, systems, exceptions and an accountable owner.

03

Evidence readiness

Representative work, outcome data and adverse scenarios can be assembled and governed.

04

Consequence and reversibility

Errors can be detected, contained, corrected and remedied within the proposed authority.

05

Platform leverage

The work strengthens reusable data, evaluation, integration, identity or observability foundations.

06

Change readiness

Operations, policy, people and support can absorb the new service without hidden manual displacement.

Portfolio rule

High value does not neutralise an unowned service, unavailable evidence or irreversible harm. Resolve the blocker, reduce the authority, redesign the service or stop the investment.

Portfolio economics

Measure the completed outcome, dependency and exit path

Token price and licence cost are inputs, not the business case. Enterprise economics include integration, human review, support, rework, control, provider concentration and the cost of changing or retiring the service.

Value

Cost per completed outcome, benefit realisation, avoided loss and service performance against baseline.

Demand

Volume, route distribution, peak capacity, queue movement and human intervention by reason.

Dependency

Model, provider, data, tool and infrastructure concentration with tested replacement paths.

Unit economics

Inference, retrieval, integration, human review, support and rework cost for one completed unit.

Reuse

Shared components used safely across workloads and the effort avoided in later delivery.

Exit readiness

Time, evidence and operational effort required to substitute, roll back, suspend or retire the service.

Published benchmark ranges can inform a hypothesis, but they do not replace a workload-specific baseline, currency, scope, measurement window and evidence owner. The Part 2 printable edition retains and classifies every figure from the earlier guide rather than presenting unscoped ranges as universal returns.

Authority context

Map obligations to the organisation, role and use case

Regulation is not one universal checklist. Identify the jurisdiction, entity, regulated role, use case, materiality and effective date before translating a source into a control or evidence requirement. The guide is an executive decision resource, not legal advice.

Source status reviewed 23 August 2026. Confirm the current official source before relying on time-sensitive material.

Primary authorities and operating references for enterprise AI implementation
AuthorityStatusApplicabilityImplementation relevance
EU Artificial Intelligence ActBinding EU lawObligations vary by regulated role, system classification and application date.Classification, transparency, records, human oversight and technical evidence.
Australian Guidance for AI AdoptionVoluntary government guidanceAustralian developers and deployers; it does not create new legal duties.Accountability, risk, stakeholder engagement, transparency, testing and monitoring.
OAIC privacy and AI guidanceRegulator guidanceInterprets Privacy Act and Australian Privacy Principles duties for entities in scope.Privacy by design, vendor diligence, personal information, retention and transparency.
APRA Letter to Industry on Artificial IntelligenceSupervisory letterAPRA-regulated entities; published 30 April 2026.Accountability, materiality, inventories, controls, third parties and operational resilience.
ASIC REP 798 and RG 271Regulatory findings and guidanceFinancial firms and complaint handling within the stated regulatory scope.Governance observations, customer impact, complaint recognition, records and escalation.
ISO/IEC 42001 and ISO/IEC 23894Voluntary international standardsManagement-system and AI risk references unless adopted by contract, policy or regulation.Management review, lifecycle risk, monitoring, corrective action and continual improvement.
NIST AI Risk Management FrameworkVoluntary US Government frameworkGlobally usable operating reference; cite the edition used.Govern, Map, Measure and Manage; evaluation and generative-AI risk treatment.
ASD secure AI development guidanceNon-binding cyber guidanceWhole-lifecycle secure design and operation guidance published by ASD's ACSC.Threat modelling, supply chain, asset records, secure defaults and incident management.
Delivery evidence

From service problem to accountable operation

Customer care shows how the four capability patterns can be composed without surrendering service ownership. Grounded self-service, governed triage, bounded agent actions and orchestration each require their own outcome, evidence and authority decision.

  1. 01

    Design

    Map demand, customer effort, complaint and escalation paths; define where a person must decide or intervene.

  2. 02

    Build

    Prove answer grounding, classification, tool access, hand-off context, recovery and service measures under representative demand.

  3. 03

    Scale

    Monitor containment with recontact, complaints, overrides, resolution and customer impact; feed failures back into evaluation.

Review first-party delivery evidence in contact centre containment, governed agent tooling and LLM evaluation.

Implementation FAQs

Implementation questions, accountable answers

Direct answers to the decisions that separate a capable prototype from a controlled enterprise service.

Put the implementation model to work

Frame the service, select the appropriate capability and establish the evidence, authority and operating model required for accountable scale.