Rafael Ferreira Souza
← All publications

Autonomy & intervention

Machines Execute. Humans Remain Accountable.

The operating policy for the Prompt Factory.

A practical model for deciding what AI may run autonomously, where people must intervene and when the system must stop.

The Factory needs a simple boundary between automation and responsibility.

Machines perform repeatable work and produce evidence. Humans define intent, accept risk, approve material changes and remain accountable for consequences.

This policy is the governance overlay for the Prompt Factory process map: the process defines where work happens; the autonomy level defines how it may run; the gate identifies who accepts the next risk.

Automation may increase. Accountability does not transfer.

01

Five levels of autonomy

Autonomy is not a single switch. It is a bounded permission assigned to a specific activity, version and operating context.

Autonomy levels A0 through A4 and their appropriate uses
LevelMeaningAppropriate use
A0Human performs the activityStrategic or exceptional decisions
A1AI assists; human completesProblem definition and architecture
A2AI prepares or executes; human approvesProduction releases and consequential actions
A3Autonomous within approved policiesTesting, monitoring, routing and reversible operations
A4Unsupervised consequential autonomyNot recommended for the first 10 clients

Three axes stay separate: the autonomy level defines how work runs; the intervention tier defines approval density by risk; the lifecycle gate defines whether evidence permits progression.

02

Process and intervention map

The Factory should operate autonomously by default, with human intervention only at decision, risk, ambiguity and accountability points.

Ten Prompt Factory stages with autonomous work, human intervention, decision gate and main KPI
StageAutonomous workHuman interventionGateMain KPI
1. Demand intakeCapture request, classify use case, detect duplicates, estimate complexity and riskProduct Owner clarifies the business outcome for new or ambiguous demandsG0 — Demand acceptedIntake lead time
2. SpecificationDraft requirements, context, constraints, test cases and success criteriaProcess Owner approves scope, expected behavior and business rulesG1 — Specification readyRequirement completeness
3. Prompt designGenerate variants, select templates, configure models, tools and contextPrompt Engineer decides architecture for complex casesG2 — Candidate readyFirst-pass success rate
4. EvaluationRun functional, regression, hallucination, latency and cost testsDomain Expert reviews subjective or business-critical outputsG3 — Quality approvedEvaluation pass rate
5. Risk and governanceScan for personal data, secrets, prohibited content, security and policy violationsSecurity, Legal or Compliance approves sensitive use casesG4 — Risk acceptedPolicy violations escaped
6. PackagingVersion prompt, attach metadata, permissions, documentation and rollback configurationRelease Owner validates audience and production impactG5 — Release authorizedRelease lead time
7. DistributionPublish to catalog, APIs, agents or workflows; perform canary rolloutHuman approval for high-impact production releasesG6 — Production activeDeployment success rate
8. MonitoringTrack quality, adoption, latency, cost, drift and incidents; trigger rollbackPrompt Ops intervenes when thresholds or SLAs are breachedG7 — Operation healthySuccess rate and cost per outcome
9. OptimizationCluster feedback, propose improvements and run controlled experimentsProduct Owner selects improvements and accepts behavior changesG8 — New version approvedValue uplift per version
10. RetirementDetect inactivity, obsolete dependencies and duplicated promptsBusiness Owner approves deprecation or replacementG9 — Retired safelyActive prompt reuse rate
The table is canonical: every stage produces evidence for its corresponding gate.
  1. 01Autonomous demand intakeClassify, deduplicate, estimate
  2. G0Demand acceptedClarify new or ambiguous outcomes
  3. 02AI-assisted specificationRequirements, constraints and tests
  4. G1Specification readyScope and business rules approved
  5. 03Prompt designVariants, models, tools and context
  6. G2Candidate readyComplex architecture reviewed
  7. 04Automated evaluationQuality, latency and cost
  8. G3Quality approvedSubjective and critical outputs reviewed
  9. 05Risk and governance scanData, security and policy
  10. G4Risk acceptedSensitive uses require approval
  11. 06Automated packagingVersion, metadata and rollback
  12. G5Release authorizedAudience and impact validated
  13. 07Controlled distributionCatalog, APIs, agents and canary
  14. G6Production activeHigh-impact releases remain human-approved
  15. 08Autonomous monitoringQuality, cost, drift and rollback
  16. G7Operation healthyPrompt Ops handles breached thresholds
  17. 09Autonomous optimizationFeedback clusters and experiments
  18. G8New version approvedMaterial behavior change accepted
  19. 10Retirement analysisInactivity, obsolescence and duplication
  20. G9Retired safelyDeprecation or replacement approved
03

Ten gates, one accountable progression

A gate is not a status meeting. It is a recorded decision over evidence. Each gate identifies the version, results, risks, accountable owner and next permitted state.

G0

Demand accepted

Is the demand real, distinct and outcome-oriented?

Accountable role: Product Owner.

Autonomous evidence

  • Request capture and classification
  • Duplicate detection
  • Complexity and risk estimate

Human intervention

  • Clarify the business outcome for new or ambiguous demands
  • Confirm accountable ownership

Main KPI Intake lead time.

G1

Specification ready

Are scope, behavior and business rules explicit?

Accountable role: Process Owner.

Autonomous evidence

  • Requirements and context
  • Constraints and test cases
  • Success criteria and operating envelope

Human intervention

  • Approve scope and expected behavior
  • Validate business rules and exceptions

Main KPI Requirement completeness.

G2

Candidate ready

Is the design coherent enough to evaluate?

Accountable role: Prompt Engineer.

Autonomous evidence

  • Prompt variants and templates
  • Model, tool and context configuration
  • Candidate package for evaluation

Human intervention

  • Choose architecture for complex cases
  • Resolve ambiguous tool or context design

Main KPI First-pass success rate.

G3

Quality approved

Does the candidate meet the quality threshold?

Accountable roles: Evaluation Engineer and Domain Expert.

Autonomous evidence

  • Functional and regression tests
  • Hallucination, latency and cost tests
  • Comparison with production baseline

Human intervention

  • Review subjective outputs
  • Decide ambiguous or business-critical cases

Main KPI Evaluation pass rate.

G4

Risk accepted

Are data, security, policy and regulatory risks acceptable?

Accountable roles: Security, Legal or Compliance.

Autonomous evidence

  • Personal-data and secret scanning
  • Security and prohibited-content checks
  • Policy-violation report

Human intervention

  • Approve sensitive use cases
  • Accept residual and regulatory risk

Main KPI Policy violations escaped.

G5

Release authorized

Is the package safe to promote for its intended audience?

Accountable role: Release Owner.

Autonomous evidence

  • Immutable version and metadata
  • Permissions and documentation
  • Rollback configuration

Human intervention

  • Validate audience
  • Approve material production impact

Main KPI Release lead time.

G6

Production active

Did controlled distribution activate successfully?

Accountable roles: Release Owner and Customer Administrator.

Autonomous evidence

  • Catalog and channel publication
  • Canary rollout and smoke tests
  • Deployment health signals

Human intervention

  • Approve high-impact production releases
  • Confirm first activation in a new context

Main KPI Deployment success rate.

G7

Operation healthy

Are quality, SLA, drift and cost within limits?

Accountable role: Prompt Ops.

Autonomous evidence

  • Quality, adoption, latency and cost monitoring
  • Drift and incident detection
  • Preapproved automatic rollback

Human intervention

  • Respond when thresholds or SLAs are breached
  • Correct, roll back, suspend or escalate

Main KPI Success rate and cost per outcome.

G8

New version approved

Does the proposed change produce meaningful value uplift?

Accountable role: Product Owner.

Autonomous evidence

  • Feedback clustering
  • Improvement proposals
  • Controlled experiment results

Human intervention

  • Select improvements
  • Accept material behavior changes

Main KPI Value uplift per version.

G9

Retired safely

Can the capability be deprecated without losing control or continuity?

Accountable role: Business Owner.

Autonomous evidence

  • Inactivity detection
  • Obsolete dependency scan
  • Duplicate and replacement analysis

Human intervention

  • Approve deprecation or replacement
  • Confirm retention of history and audit evidence

Main KPI Active prompt reuse rate.

04

Intervention policy by risk

Autonomy changes with consequence. The same technical workflow can operate under a different approval pattern when data, regulation, reversibility or customer impact changes.

Green

Low risk

Fully autonomous flow; humans audit a sample.

Yellow

Medium risk

Human approval at specification and release.

Red

High risk

Domain, security and release approval required; no autonomous production deployment.

Critical

Stop first

Automatic suspension or rollback, followed by mandatory incident review.

Humans intervene whenever there is ambiguity, sensitive data, regulatory exposure, irreversible action, quality below threshold, unusual cost, model drift or material customer impact.

05

Runtime intervention rules

Factory gates control how a capability is produced. This matrix controls what the capability may do after deployment.

Runtime actions and their default autonomy
Runtime actionDefault autonomy
Search approved internal knowledgeAutonomous
Summarize or classify informationAutonomous with monitoring
Draft emails, reports or proposalsAutonomous draft; human decides use
Recommend a decisionAutonomous recommendation; human accountable
Update a reversible CRM fieldAutonomous if preapproved and audited
Send external communicationHuman approval initially
Publish public contentHuman approval
Purchase, payment or financial commitmentHuman approval
Delete recordsHuman approval
Change permissions or credentialsHuman approval
Sign or accept contractual termsHuman only
Health, employment or credit decisionHuman-controlled high-risk process
Switch to an already-certified modelAutonomous
Introduce a new model or toolHuman approval and reevaluation
06

Mandatory escalation triggers

The runtime must stop or request human review whenever an execution leaves its approved purpose, evidence or operating envelope.

  • Input is outside the approved purpose
  • Sensitive data is detected unexpectedly
  • An evaluation or confidence rule fails
  • Model providers disagree materially
  • A tool requests broader permissions
  • An action is irreversible
  • Cost exceeds the tenant budget
  • Quality declines after a model change
  • A user reports harm or serious error
  • Cross-tenant access is attempted
  • An unapproved model, tool or data source is requested
  • Regulatory exposure is detected
  • Model drift changes expected behavior
  • Material customer impact is possible
For critical security events, the platform should block first and investigate second.
07

Policy for the first 10 clients

Early operation should favor evidence over speed. The autonomy envelope can expand only after the Factory has learned from real performance.

Conservative autonomy posture

  • New or ambiguous demands receive Product Owner review at G0
  • Medium-risk flows require human approval at specification and release
  • High-risk flows require domain, security and release approval, with no autonomous production activation
  • Autonomous prompt design remains inside the approved operating envelope
  • Automated evaluation with mandatory human sample review
  • First production batch reviewed intensively
  • Sampling reduced only after sufficient performance evidence
  • No autonomous introduction of models, tools or data sources
  • Automatic suspension or rollback when critical thresholds fail
08

The accountability rule

As evidence accumulates, routine and reversible operations can move from A2 to A3. Purpose, risk acceptance, material scope changes and consequential actions stay attached to named people.

This is how the Factory gains speed without confusing execution with authority.

Machines can prepare, test, monitor, route and roll back. Humans remain responsible for why the capability exists, where it may operate and what consequences the organization accepts.

The complete system

Principles define the promise. Processes organize the work. Governance controls autonomy.