Structure B, adopted by the working group in cycle 1 (vote 6 · 4 · 1; dissent recorded below): records and accountability are separate principles, the former principle 1 became the test of purpose in A2, and cycle 1 closed the set at seven — and v1.0 ships seven. A candidate eighth principle stood in the release candidate; the cycle-3 coverage question put it to the working group, and the project chair withdrew it at the freeze. The vote, the candidate's full text and the withdrawal are recorded in The eighth-principle question at the end of this section. Each principle carries its statement, what the community's data says, a practice note, its minimum viable version, and its status. A cross-walk from these seven back to the webinars and the earlier numbering is in D4.
P1 — AI amplifies data; it does not fix it or guarantee its quality (cycle-1 P2)
AI makes the data you already hold more influential and more confident-looking. It does not make it more true.
Data quality, provenance and the Verified / Inferred / Unknown distinction are governance preconditions for trusting any AI output. Before any AI use, name the system of record for each source the AI reads; where two authoritative sources disagree, record which one governs. A value labelled Inferred carries the evidence and reasoning used to infer it — a citation a human cannot check is not provenance. Every AI output carries the data-quality status and key uncertainties of its sources; higher-consequence decisions require explicit verification before the output is used. AI must make uncertainty visible, not hide it behind confidence. No AI-authored value enters an authoritative register without its label and a person's sign-off.
What the community's data says. Rated 3.86/5 (n=71; 70% rated 4–5) [W1 q9]. 75% have seen bad data become a confident output; 45% the authority illusion; 28% silent fabrication; 25% data laundering [W1 q8-patterns]. The whole series' top-upvoted worry: "generating bad data and results" [W1 q2]. In cycle 1 this principle drew five break-cases and all three reviewer responses — the most load-bearing principle in the set; its problem was sufficiency, not wording. The strongest case against "AI will clean it for us" as a management belief came from the working group itself.
In practice. Proportionate sign-off: class or batch sign-off with sample checks for routine fields; per-value sign-off for critical fields. Labels attach to owners, not processes (C1). Non-authoritative fields may be cleansed at level 4 within an envelope; authoritative records cap at level 3 (B3 rows 1–2).
Minimum viable version. Default every unlabelled legacy value to Unknown, then label on touch.
Status: Closed (consensus), cycle 1.
P2 — An AI prediction is not a risk decision (cycle-1 P3)
AI can estimate condition, failure probability or risk. Before a prediction triggers action, a competent reviewer validates calibration, data quality and drift (is the asset still in-distribution?), and the accountable person accepts and records the residual risk. AI may set out the residual-risk options; only a person accepts one. A prediction lacking a recorded validation and a recorded risk acceptance is deferred, not actioned. Accountability never transfers to the model.
What the community's data says. Rated 3.76/5 (n=46) [W2 principle-test]; on acting on a prediction the room capped AI at analyse-or-recommend (67%, n=42) with autonomy at 5%. Asked who owns the risk when AI flags, answers converged on a named human — the asset owner, the asset manager, "whoever normally owns the risk, regardless of AI involvement." In cycle 1 this was the most reasoned support in the set, with zero break-cases. One working-group note is the best one-line case for calibration validation: the AI may not have any examples for large-consequence failures.
In practice. Proportionate validation — depth scales with consequence; a small organisation follows a defined minimum validation path rather than a full calibration-and-drift programme. Prediction-triggered work carries a risk-acceptance record: who accepted, on what evidence, with what reconsideration date.
Minimum viable version. For one prediction-driven workflow, add two fields to the work order: validated by and residual risk accepted by. Blank means deferred.
Status: Closed (consensus), cycle 1.
P3 — Humans decide what "better" means (cycle-1 P4)
AI can search for the best answer. It cannot decide what "best" should mean.
The decision owner — the accountable person for this decision — sets and re-signs what the organisation is optimising for: the asset-management objectives being served, the trade-offs between them, the limits any answer must respect, and the method used to find it. No AI output may set or silently change any of these. The objectives and weights are re-signed whenever they change and at least on the interval the organisation sets for the envelope or plan they govern (B5); an unsigned or expired set is a draft. Where an optimiser proposes deferring work, each deferral is recorded as a named, dated risk acceptance (C1, the deferral ledger).
The principle applies to every kind of optimisation — capital programmes, maintenance schedules, inventory, dispatch — and to the weights that arrive inside a product: a default nobody signed is still a value choice, made by the vendor.
What the community's data says. The signed objective rated 3.73 (n=33); the deferral ledger 4.07 (n=30 — W3's best); the frontier right 3.39 (n=36 — the series' weakest, and now a practice note). Behaviour was sharper than the ratings: 85% would not have signed the vendor-default weights nobody in the room chose; 64% rejected the optimiser's silent trade; in only 39% of organisations could a named person show the value weights behind the capital plan today (n=36) [W3]. What the room valued when it held the pen: cost 26 · risk 25 · service 24 · equity 15 · carbon 10 (medians, n=45).
In practice. The value model is an owned, signed, re-signed artefact — like accounts. Showing the frontier: where the tooling allows, the people affected by a plan see the near-optimal alternatives and what each trades away, and choose among them; where it does not, the owner at least records which alternatives were considered. The cross-department check — who owns finding out what the optimiser quietly defunded — is a named role in the matrix's accountability column (B3 row 6).
Minimum viable version. If nobody has signed the objectives and weights, the output is a draft and does not proceed to approval.
Status: Closed (consensus), cycle 3 — the wording was rewritten twice in cycle 1; the re-presented text passed the last call 9–0 with no objection.
P4 — AI may advise widely; it may act only narrowly (cycle-1 P5)
Automation depends on consequence, reversibility and oversight.
Human approval is the default for any action an AI initiates. Where an AI is allowed to act, the action is limited in scope and volume, reversible within a stated window, under oversight with enough time to intervene, and clearly owned — with a stop that has been engineered and practised, not assumed. High-consequence or irreversible actions stay under human control, and every organisation names, in its own matrix, the actions that must always be made by people — safety-critical and statutory work orders among them.
The envelope's anatomy — eleven parts, from scope and magnitude to the warn band and edge behaviour — lives in B5, not here.
What the community's data says. W4's paired experiment: after watching a bounded agent act for forty minutes the room moved +0.44 rungs (paired, n=25) — to act with approval (26% → 60%), not autonomy (13% → 9%). The reversibility premium: 38 points — the same work order drew 59% "allow auto" when cancellable for two hours and 21% when cost committed on creation. The room's first volume budget: median 7.5 machine-created work orders per planner per week before auto-pause, 30% never. Trigger gate rated 4.15 (n=33), engineered stop 4.06 (n=35). Stop authority was the only retained right nobody dropped after the storm (48% → 64% paired). "Never" appeared only when the case was concrete — 19% on a live 66 kV isolation, 0% in every abstract vote [W4 · SES F1–F3, F5].
In practice. Every acting AI has an envelope document and a named owner; the stop is a tested control, drilled like an emergency procedure, not a menu item. Monitoring that arrives after the action is not oversight.
Minimum viable version. Name one class of action AI may take, one volume limit, one accountable person, one tested stop.
Status: Closed (consensus), cycle 1.
P5 — No AI action stands without a record (cycle-1 P6, first half)
AI may act or defer; nothing counts until it is logged. A person must be able to reconstruct what was done, what was declined, and who owned the call, from a durable record the organisation controls.
A record must be able to show what was not done. An approval states what was examined and what was accepted unexamined, measured against a review capacity declared in advance. Anything beyond that declared capacity is deferred, not approved — an approval signature that looks identical whether the person examined three items or forty is the failure this principle exists to prevent.
Every AI use carries a live register entry with six core fields: named owner, approval status, operating envelope, validation record, review frequency, retirement trigger (the full schema, which extends these six, is in C1).
What the community's data says. The ledgers won their rooms: the action ledger is the series' highest-rated principle (4.27, n=33; 82% rated 4–5) and the deferral ledger W3's highest (4.07, n=30; 83%); 0% believed a re-signing schedule would be the safeguard their organisation quietly skipped; the community consistently endorsed recording mechanisms above access rights and process promises [SES F4 · W4 §9]. The principle's opening and its second paragraph were written by working-group members in cycle 1 — the strongest single argument for Structure B came from the volunteers, not the drafters.
In practice. Records are the artefact that makes oversight verifiable (P6) and the loop closable (C1, C3): predicted-versus-realised is the one record a rubber stamp cannot produce.
Minimum viable version. For one AI use already in service, write its register entry with all six core fields. A field you cannot fill is a finding.
Status: Closed (consensus), cycle 1.
P6 — Accountability is named, non-delegable, and carries the power to halt (cycle-1 P6, second half)
"A human reviews it" is a claim, not a control. Oversight counts only when it is verifiable: interrogable outputs, visible limitations, and effective challenge by people with the capability, the authority and the time to say no.
Every AI use has one accountable person with the authority to halt it. Competence is a condition of that role, not an assumption — the accountable person must be able to detect that the output is wrong, and the required competences are scheduled in C3. No person challenges their own work: independent challenge applies to high-consequence AI use; for other uses the accountable person verifies, and accountability remains theirs either way. "The model said so" is no defence.
The accountable person, defined. The named individual who holds the authority to approve, challenge and halt a given AI use and who answers for the decision it changes. Concretely, in an organisation's own terms:
- Who: one person per AI use (or per matrix row), named in the register — a role-holder, not a committee, and not the vendor. Where the organisation uses ISO 55001 delegations or ISO 19650 party roles, the accountable person is whoever already holds the delegation for that decision class; the framework does not create a new office.
- Authority: can stop the AI use, no reason required; can reject or amend any output; can change its level in the matrix; signs the objectives and weights where it optimises (P3) and the envelope where it acts (B5).
- Competence: is able to tell when the output is wrong, or has access — resourced and recorded — to someone who can. The accountable person is not necessarily the competent reviewer (P2) and in high-consequence use must not be (independent challenge); but they are responsible for ensuring the competence exists before relying on the output (P7). Whether they are also the competent person is determined by the competence schedule in C3, recorded in the register.
- Cannot delegate: the authority can be exercised through others; the accountability cannot be handed to a model, a vendor, a committee or a subordinate.
What the community's data says. Confidence in declared oversight is low: median guess for the share of AI-generated plans audited today — 0% (n=35) [W3]. 14% have already seen rubber-stamp oversight; 45% the authority illusion [W1 q8]. Asked which safeguard their organisation would quietly skip first, the close-out feedback loop (30%) and the practised stop drill (22%) topped the named list [W4 §9]. Under pressure the room's approval and stop rights rose (paired +12pp, +16pp) while abstract stewardship fell [W4 §1e · SES F2, F8]. In cycle 1, eight comment instances from six working-group members asked for one concrete term — this is it.
In practice. For high-consequence decisions: challenge is resourced and recorded; oversight quality is audited (spot-checks of approvals), not assumed; every AI use case has one named accountable person who can halt it and whose name is in the matrix.
Minimum viable version. Write one name against each AI product already in service. "Nobody" is a finding.
Status: Closed (consensus), cycle 1 — the definition above answers the 14 Aug call's decision D2 ([C1-log]) that it be concrete and inside the framework.
P7 — AI must not erode the expertise the framework depends on (cycle-1 P7)
Every principle above assumes a person who can still tell when the AI is wrong.
AI use must therefore preserve that capability: the reasoning that supports an answer travels with the answer — both are required; competence in the underlying judgement checked and recorded for the people who approve AI outputs; and over-reliance actively watched for. Where the capability no longer exists, the AI use is not approved — it is deferred, with the capability gap escalated to someone who can make the business decision: fill the gap, or consciously accept indefinite deferral. That person is not necessarily whoever previously owned the AI activity. The framework sets no duration for the escalation — it depends on each organisation's policy, budget and risk appetite — but it requires that the deferral has an owner and a decision, so that a register of deferred uses does not grow because nobody decides.
What the community's data says. This remains the least-quantified principle — no boundary vote tested it; the signal is in open text ("loss of abilities and skills to solve problems", "cognitive offloading", "loss of ability to evaluate AI responses") [W1 q2, q10]. The working group kept it as a principle in cycle 1 because it is the precondition for P2, P4 and P6, made its language mandatory ("should augment" → "must not erode"), replaced rotation through manual practice with a competence check (many teams have one expert), and supplied an evidence design: survey the community on observed skill loss now, and in a later cycle ask who runs competence checks and whether scores held (D1, OQ-05). The case for keeping it came from the framework's own adversarial test: the reviewer scenario's third failure happened because the one engineer who remembered why the main mattered had retired.
In practice. Capability impact is assessed at screening (B4 q6); augmentation patterns (AI drafts, human decides; AI explains, human verifies) are preferred over replacement patterns for judgement-bearing tasks; the competence schedule and tacit-knowledge practice are in C3.
Minimum viable version. For one AI use, name the person who could tell if its output were wrong. If the answer is "nobody any more", defer the use and escalate.
Status: Closed (consensus), cycle 1; this text carries the three wording changes the 14 Aug call requested, and one freeze-pass grammatical repair logged in D4.
The eighth-principle question
Cycle 1 closed the principle set at seven. Two candidates for an eighth stood at the final pass, and the cycle-3 coverage question (CQ) put both to the working group directly — a named vote beside the last call, because both candidates had reached the release candidate without a group decision behind them.
The two candidates.
- Close the loop — AI is judged by realised outcomes, not outputs — the project chair's candidate (drafted 16 Aug from the cycle-1 and cycle-2 evidence; added to the release candidate as P8 on 23 Aug, flagged as not yet having been before the group). Its full text as it stood in the release candidate is preserved below.
- Affected parties — a principle for the people, groups and communities on whom an AI-influenced decision lands, especially when the output is wrong. Proposed by a working-group member in cycle 1 and raised twice without resolution on a call; the release candidate carried it not as a principle but as a required register entry (who is affected, how they are told — C1) and a challenge path (C3).
The vote (nine of record, 24–27 Aug 2026; every vote carried a written reason):
| Option | Votes |
|---|
| (a) Keep P8 — eight principles; affected parties stay with the register entry and challenge path, revisited after v1.0 | 5 |
| (b) Keep P8 and add affected parties — nine principles | 1 |
| (c) Drop P8 — back to the seven agreed in cycle 1; the duty to check outcomes stays as a records requirement with a named owner and a date | 2 |
| (d) Drop P8, but add affected parties — eight principles, the other one | 1 |
Keep the eighth principle: 6 – 3. Affected parties as a principle of its own: 2 – 7. The full record — every vote with its reason, by name — is in the cycle-3 disposition log (D4).
The project chair's decision — recorded here so it cannot be mistaken for the group's. The vote favoured keeping the eighth principle, and option (a) alone was an outright majority. The project chair nonetheless withdrew his own candidate at the freeze, and v1.0 ships seven principles. The dissent persuaded him: the strongest reasoned argument on the record holds that verifying outcomes and learning from experience — important as it is — is a mandatory governance and assurance requirement that demonstrates compliance with P1, not an additional principle; a second dissent was not convinced an eighth principle is needed, holding that the wider circle of interested parties is already party to P1–P7; a third would drop it and strengthen P4's target-setting instead. Against that stood a candidate the project chair had added himself on 23 Aug, never proposed in any working-group response, on its first and only review. A project chair's addition carried over three members' recorded dissent would stand differently in this set than the seven principles the group proposed, contested and closed together — and deference to the dissent costs the framework nothing it cannot afford, because nothing of the duty is lost: the close-the-loop requirement — the outcome record, its owner, its date, its two feedback directions — is mandatory in C1 and C3 exactly as the drop option described, and it binds at ladder levels 3 and 4 (B1).
The keep reasons, answered. The six keep votes were reasoned, and the strongest deserves its answer here, not only in a log. One keep reason warned that leaving the duty in the records section "buries the most-skipped control in the place things get skipped from". The withdrawal does not do that: the outcome record is not prose in a records section — it is a mandatory register field with a named owner and a date, and it fails closed: a use whose outcome field stays empty is deferred at its next re-sign, not renewed (C1). What was withdrawn is the label, not the gate. A second keep reason held that asset management is ultimately judged by the outcomes delivered over the asset lifecycle — that case travels with OQ-19 back to the working group, which decides the principle question in v1.x on a group vote. Every coverage-question reason, keep and drop, is preserved by name in the cycle-3 record (D4).
What reopens it. Whether close-the-loop deserves principle status returns to the working group in a v1.x cycle with this record attached (D1, OQ-19) — where, if adopted, it will carry a group vote as its provenance rather than a project chair's addition. The affected-parties question stays open the same way (OQ-20), with the majority option's own wording — revisited after v1.0 — as the standing commitment.
The candidate text, as it stood in the release candidate (preserved for the v1.x decision; not a principle of v1.0):
Close the loop: AI is judged by realised outcomes, not outputs. Every principle above governs the moment before or during an AI-influenced decision. This one governs afterwards. Every AI use is judged by realised outcomes, not outputs. On the schedule in the register, the accountable person compares what the AI claimed with what happened — the deferred mains that failed, the predicted failures that did not, the work orders that were cancelled. Misses feed back in both directions: into the governance side — the envelope, the approval status, the validation record — and into the AI itself: the same outcome record that audits the model is the evidence that recalibrates, retunes and retrains it. Governance and improvement are one loop, not two systems. An AI use that generates no outcome evidence has no evidence that it works — and a use whose outcome field stays empty is deferred at its next re-sign, not renewed.
What the community's data says. The community's median guess for the share of AI-generated plans audited today is 0% (n=35) [W3 w3-p2-audit-guess] — the one control it could not estimate above zero. Asked which safeguard their organisation would quietly skip first, the close-out feedback loop topped the named list at 30% [W4 §9]. In cycle 2 the boundary test's q9 — will anyone find out whether the deferred mains failed? — drew a near-unanimous no (1 yes · 10 no); one member's reason names the failure mode: a wrong deferral surfaces as a main failing in the ground, not as a detected error. A cycle-1 note supplies the other half: the AI may hold no examples of large-consequence failures — the calibration P2 demands in year two is only possible if someone kept score in year one. No other principle looks back: P2 validates before action, P5 records what was decided, P6 oversees at the moment of approval. The candidate passes the framework's two design rules natively — predicted-beside-realised is the one artefact a rubber stamp cannot produce, and because the outcome record is also what improves the model, keeping it is productive work rather than compliance overhead.
Minimum viable version. For one AI use, record prediction beside outcome for its next ten decisions; review once, and act on what you find — on the envelope, or on the model.
The other candidates, for the record. Six of ten working-group members proposed an eighth principle in cycle 1; every proposal was adopted somewhere — two became P5 verbatim in substance, one became P1's spine, one the test of purpose's capability condition, one the C3 competence schedule, and the affected-parties proposal became the C1/C3 practice requirement above. The project chair contributed two further candidates with AI-assisted provenance: a default is a decision is carried as a clause in the test of purpose (A2) and in C4, and the cheap-path rule as a design rule in A1 — deliberately not principles.
Provenance. Drafted by AI (Stream 1) from: [W5B] §3 · [SES] F2, F4, F6–F8 · [W1]–[W4] as cited. Reviewed: WG cycle 1 (MA1: 11 responses, no principle majority-rejected; MA2: 11 break-cases; MA3: 10; reviewer packet 1: 3) with every comment answered in [C1-log]; structure vote B 6 · A 4 · 1 abstention; the 14 Aug call's decisions D1–D6 and the HANDOFF §2 rewrites applied here. Cycle-3 last call: 9 yes · 0 objections; the coverage question voted 6–3 to keep the candidate eighth principle and 2–7 against an affected-parties principle; the project chair withdrew his candidate at the freeze — the decision and its reasons are recorded in this section, and the minority and majority reasons are in [C3-log]. Consensus: rough consensus declared by the project chair, 14 Aug 2026 (the seven principles) · Dissent recorded: four members voted for Structure A; the instrument gave them no reason field, the log says so and attributes no rationale to them; the D4 cross-walk answers the traceability case at no cost. One member's structure and ladder responses are recorded as abstentions after a readability failure the draft owned — the plain-language pass and the ladder's printed one-line definitions are its fix. Signed off: the project chair, 31 Aug 2026. Status: Frozen v1.0. Last updated: 31 Aug 2026.