Asset Zero · Progress & telemetry

Watch the project run itself

To our knowledge, the first project to pre-register and publish per-function human-AI boundary levels with enforcement telemetry. Every AI action — and every denial — lands in the public ledger below. The levels themselves are on the about page.

📡 RSS — status changes, not prose

The AI action ledger

Append-only and hash-chained; written by enforcement hooks, not by the AI; denials are first-class entries. Exported to the repository nightly once the chain is live — the ledger is the product.

Loading the ledger…

The co-production cycles

The framework closes section by section through August — the calendar, the weekly rhythm, the hours ask and the watch metrics, published as they are. Sections close on the living framework.

BuildMon 3 – Fri 7 Aug✓ done

Stream 1 drafts the complete framework v0.5 (all 18 sections, provenance-tagged, from the evidence spine + W5B) · fact-check · Kai signs · kickoff email goes Mon 3 or Tue 4 Aug · volunteers confirm role + the August-hours consent (§4)

Closes:

v0.5 full draft

Kickoff + Cycle 1Mon 10 – Fri 14 Aug✓ done

Kickoff call Mon 10, 12:00–12:30 [TBC] · cycle-1 packet lands same day · micro-asks close Wed 12 noon · call Thu 13, 12:00–13:00 · dispositions + close Fri 14 · Reviewer packet 1 (Mon 10 → Fri 14)

Closes:

A2 + B2 + B1 — one finding, seven principles + structure question, ladder & human roles → v0.6

Cycle 2Mon 17 – Fri 21 Aug✓ done

Packet Mon 17 · asks close Wed 19 noon · call Thu 20, 12:00–13:00 · dispositions + close Fri 21 · Reviewer packet 2 (Mon 17 → Fri 21)

Closes:

B3 + B4 + B5 — Decision Rights Matrix, boundary test, envelopes/trigger/stop → v0.7

Cycle 3 — final passMon 24 – Fri 28 Aug

Whole-document last call (async, IETF-style): 72-h objection window Mon 24 – Wed 26 on the full v0.7 + first-look C/D/A sections · recommendations priority vote (form, 15 min) · optional drop-in clinic Thu 27, 12:00 (attendance not required, not counted in anyone's budget) · dissent recorded · chair closes all sections Fri 28

Closes:

C1–C5 + D1–D4 + A1/A3/A4 + whole-doc → v1.0-rc

FreezeMon 31 Aug

v1.0 tagged · snapshot PDF assembled and published · participation stats + disposition archive into D4 · thank-you + credit confirmation to every volunteer

Closes:

v1.0

The one-week cycle rhythm
DayWhatWhoEffort
MonPacket out (brief + drafts + asks + reviewer case)AI drafts · Kai approves & sends
Mon–Wed noonMicro-asks (2–3 structured, ≤40 min total)Working Group≤40 min
Wed noonResponse check: <60% → Kai's personal two-line nudge same afternoonKai
Thu 12:00The call — contested points only, 60 min, rough-consensus testWG + chair1 h
Mon–FriReviewer packet (one case + 3 questions)Reviewers≤1 h
FriDisposition log closed (AI drafts, chair adjudicates every row) · section status set · revised draft + log publishedChair + Kai

Missed a call? Recorded + 15-min catch-up pack; async input carries identical weight — decisions close on the log, not in the room. Rough consensus needs objections addressed, not attendance; participation stats publish transparently in D4 either way.

The honest hours ask

The signup promise was ≤4 h/month until October — a ~10–12 h total journey. The compressed ask is less total, more concentrated:

  • Working Group: budget 5 hours, once — the scheduled components sum to ≈4 h 35 min at their maxima (kickoff 30 min + 2 × (asks ≤40 min + call 1 h) + final-pass read & vote ≤45 min); we ask people to budget 5 for reading time. That is 35–60 min over the advertised August cap, once — and then the project is done two months early.
  • Reviewers: ~2 hours — two 1-h packets; the final-pass read is optional.

The kickoff email asks each volunteer to explicitly confirm they can wear ~4½–5 h (WG) / ~2 h (Rev) in August — informed consent, in writing, per the project's own ethos. Anyone who can't still counts and is still credited: async-only participation on whatever they can give.

Watch metrics & risk plan — published, not private
WatchThresholdResponse
Micro-ask response<60% by Wed noonPersonal nudge same day → if still short Fri, section closes on received input with participation stats published; last call catches stragglers
Call attendance<8Proceed; decisions close on the log; catch-up pack
Reviewer returns<6 of 10Halve packet 2; rotate cases toward confirmed reviewers
Consensus stallAny section contested past its FridayChair closes with dissent recorded — no section gets a second week
Volunteer overload signalsAnyDrop to 1 micro-ask; the ~5 h consent is a ceiling, not a target
Build week slipsv0.5 not signed by Fri 7Cycle 1 runs on B2-A2 v0.2 (already built + fact-checked); B1 joins cycle 2
Tool (Armature)Not readyEmail + plain forms — zero process change (this was always the fallback)

Cycle 2 packet — what volunteers are working with

Published in full, per the build-in-public rule; the calendar above carries the dates. The disposition log publishes with the closed sections.

B3 · B4 · B5 — The Instruments (Draft v0.1)

B3 · B4 · B5 — The Instruments (Draft v0.1)

Stream 1 · Asset Zero. Drafted 17 Aug 2026 for cycle 2 from the W5 synthesis brief (§4 draft matrix, §5 boundary test) and the Series Evidence Synthesis (Findings 1, 3, 5, 6, 8). Every level, question and number below is a draft for the Working Group to judge — the matrix rows are hypotheses populated from the community's boundary votes, not asserted answers. Structure B numbering; "accountable person" carries both accountability and capability (14 Aug call, decision D2), with its concrete definition still owed. AI-drafted; Kai approves; the Working Group decides.

Section IDs here are framework sections (Part B — The boundary): B3 the Decision Rights Matrix, B4 the boundary test, B5 envelopes, triggers and the stop. Where a principle is meant it is written P1–P7.


B3 — The Decision Rights Matrix

What it is. One table an organisation fills in for itself: for each class of asset-management decision, the maximum AI role level (the series ladder: 0 None · 1 Assistant · 2 Analyst · 3 Recommender · 4 Controlled actor · 5 Autonomous), what humans must retain regardless of level, and the accountable person by name. Levels 4 and 5 always mean within a pre-approved, bounded, reversible envelope with an engineered stop (B5) — never autonomous on safety-critical or irreversible decisions. The rows are written as concrete work classes, because across every abstract boundary vote in the series "should not be used" took 0%, while a live 66 kV isolation put in front of the same community drew 19% "never" (Finding 3).

How to read the drafted level. Each row carries the level the series' own votes support. Where the community's mode was split between two levels, the row is drafted at the level the reversibility evidence supports and the split is stated. The W5 brief's v0.1 hypothesis is noted where this draft departs from it. Levels 4 in this table are always "act within limits" — the limit is the envelope, and the envelope is B5's job.

#Row idAsset decisionDrafted AI levelHumans must retainWhy this level (the evidence)
1data-cleanseAsset-data extraction and cleansing — non-authoritative fields4Data authority; the verified / inferred / unknown label on every field the AI touched; the accountable person for the registerW1: 68% capped AI at analyse-or-recommend on facts about the asset (n=78), but that vote covered authoritative facts. For non-authoritative fields the act is reversible and labelled — hence 4, with provenance retained. The row most likely to be contested; propose 3 if you read W1 as covering all fields
2auth-recordsChanges to authoritative asset records3Sign-off on critical fields; provenance; who owns the recordW1 q5: 68% analyse-or-recommend cap on facts about the asset; W1 governs-the-knowledge was a top retained right (66%). An authoritative record is the thing everything else trusts — the AI may propose the change, a person commits it
3condition-classCondition classification from imagery and sensors3Validation of unusual and high-risk cases; the calibration recordW2 dr-poll-c_cond (n=47): 38% Analyse, 38% Recommend, 11% Act with sign-off, 0% autonomous. Drafted at the upper mode (3) because classification is reversible and reviewable in bulk; the human retains the outlier check
4health-predictionAsset-health and failure prediction3Calibration check; residual-risk acceptance stays human (P3)W2 dr-poll-c_perf (n=44): 57% Recommend, 7% Act with sign-off, 2% autonomous. A prediction is not a risk decision — P3 keeps acceptance human, so 3 is the ceiling
5criticality-riskAsset criticality and risk scoring2Risk appetite; consequence valuation; the weights behind the scoreW2 dr-poll-c_risk (n=46): 50% Analyse, 37% Recommend, 4% Act with sign-off. The community's mode is Analyse. W5 v0.1 hypothesised 2–3; drafted at 2 because the score encodes value judgements (Finding 6) — propose 3 if your organisation signs the weights explicitly
6capital-prioritisationRenewal and capital prioritisation3The value framework and its weights, signed by the accountable person; the trade-off; the choice; what the optimiser quietly defundedW3: 72% capped AI at analyse-or-recommend on money and 0% chose autonomy (n=18); 85% would not have signed vendor-default weights (n=39); shown what an optimiser quietly defunded, 64% rejected the trade (n=42). The AI may rank; the human owns "better" (P4)
7spend-within-capCommitting spend within a pre-set cap — ordering parts, hiring plant3Rule-setting; the cap; exception approval; the reversibility windowW4 C18 (n=34): 35% Recommend, 26% Act with sign-off, 6% autonomous — two-thirds at 3 or below. Cost that commits on creation is the least reversible act in the series (the 38-point reversibility premium, Finding 3). W5 v0.1 hypothesised 4; drafted at 3, and 4 is defensible only with a hard cap and a cancellation window — say which
8risk-acceptanceRisk acceptance and service-level change2Accountability; stakeholder and executive judgement; the recorded acceptanceW2 dr-poll-c_accept (n=41): 44% Analyse, 29% Recommend, 7% Act with sign-off. Accepting a risk is the decision the whole framework exists to keep human (P3, P6); the AI informs it
9routine-work-ordersRoutine, low-risk, reversible work orders — creation and triggering4The definition of "low-risk"; the volume cap (B5); the reversibility window; stop authority by nameW4 envelope votes: 42–53% allow auto-trigger for five reversible work classes, 0% never (n=31–36); the same pothole inspection drew 59% allow when cancellable for two hours vs 21% when cost commits (n=29/33). The warm boundary vote moved the room to 60% "act with human approval" (n=35). Level 4 is earned by the envelope, not assumed
10crew-dispatchMaintenance scheduling and crew dispatch — booking people against work4Operational feasibility; override; the roster ownerW4 C18 (n=34): 41% Recommend, 35% Act with sign-off — the highest act-with-sign-off share of any decision-rights poll in the series. Scheduling is reversible and continuously overridden in practice; drafted at 4 within limits, propose 3 if dispatch commits crews to hazardous work
11spares-replenishmentLow-value spares replenishment4Criticality tiering; the value threshold; spend oversightW4 E5 spares reorder below $500 (n=36): 47% allow auto-trigger, 50% approval gate, 3% never — the only reversible class where a "never" appeared, because spend commits. Drafted at 4 under a value cap and a criticality tier the human sets
12safety-criticalSafety-critical intervention and control actions1Final authority; risk ownership; an independent safety layer; the concrete "never" listW4 E6 live 66 kV feeder isolation + permit to work (n=36): 11% allow, 69% approval gate, 19% never — the series' one concrete ban. W5 v0.1 hypothesised 1–2; drafted at 1 (assist only) so that any organisation choosing 2 does so deliberately and writes down why

The named-accountability column. Every row in an organisation's own matrix carries the accountable person by name — the person with the power to halt (B4 Q7). It is left blank here on purpose: it cannot be drafted by anyone but the organisation.

Minimum viable version. Fill in three rows — the one you already automate, the one you are being sold, and the one you would never automate — and name the accountable person for each.

Provenance: [W5B] §4 · Synthesis F1, F3, F6, F8 · [W2] dr-poll-c_* · [W4] §6–7 · [W1] q5-ai-role, q6-humans-retain · [W3] §1a, w3-s1-r5-sign-default, w3-s2-r5-accept-trade.


B4 — The boundary test

What it is. Nine questions asked, in order, before an organisation relies on AI in an asset-management decision — the screening tool that decides whether and at what level a decision enters the matrix. Reordered from the W5 brief so that the test opens with the decision, not the technology, and asks early whether the decision should be automated at all.

The scenario

A water utility's planning team runs a vendor optimiser over its sewer-main renewal programme — about four thousand mains. The tool proposes deferring roughly a fifth of the planned renewals by five years to fit the capital envelope, ranking mains by predicted failure probability. The value weights behind the ranking shipped with the product; nobody in the utility signed them. Around a third of the condition records the model reads are inferred rather than verified — the tool does not show which. The head of planning wants to adopt the plan as the year's programme and have the tool write the deferral entries straight into the register so the team is not retyping four hundred lines. Nothing has gone wrong. Yet.

(A composite built from the series' own patterns — vendor-default weights nobody signed [W3, 85%], weights a named person could show today [W3, 39%], the optimiser's quiet trade [W3, 64% rejected], inferred data presented as verified [W1, W2 cases]. No real utility is described.)

The nine questions

  1. Is it clear which asset-management decision is being changed — the five-year deferral of specific renewals — rather than "what the optimiser can do"?
  2. Should this decision be automated at all — is a five-year deferral of a fifth of the programme a decision an organisation should ever let a tool commit without a person choosing it?
  3. Is the AI's contribution correctly placed at Recommend here — or, by writing the deferrals into the register, is it in effect acting?
  4. Is the consequence if the ranking is wrong tolerable, and is it reversible within a stated window — can a deferred main be pulled back into the programme once the money is committed elsewhere?
  5. Is the evidence and validation required to rely on this output in place — provenance labels on the inferred condition records, a calibration record for the failure model, an in-distribution check?
  6. Can the person validating the output — the planner reviewing four hundred lines — meaningfully challenge it: the capability, the authority, the time, and the evidence to do so?
  7. Is it clear who approves the programme, and who can override or stop the register write — and how fast?
  8. Is there an accountable person, by name, for the deferral decision — someone with the power to halt it, who has also signed the value weights?
  9. Will errors, drift or unintended consequences be detected over time — is anyone scheduled to find out whether the deferred mains failed?

How an organisation uses it. Any "no" on questions 1, 2, 7 or 8 stops the AI use before the matrix is consulted. A "depends" is a finding: it names the condition the envelope (B5) must carry. Scoring and a criticality-keyed screening guide follow in the practice notes (C3/C5), not in the test itself.

Minimum viable version. Ask questions 1, 2 and 8 of one AI use already in service. If any answer is "no", that is the first thing to fix.

Provenance: [W5B] §5 (nine questions; reordered per Framework-v0.1-Skeleton §2 B4) · Synthesis F6, F8 · [W3] w3-s1-r5-sign-default, w3-s1-r6-weights-written, w3-s2-r5-accept-trade · [W1] q8-patterns.


B5 — Envelopes, triggers and the stop

What it is. The numbers that keep any level-4 action inside human-set bounds. Where B3 says an AI may act within limits, B5 is where the limits are written down. The community is not pricing "AI acting" in the abstract — it is pricing reversibility (Finding 3), so the envelope is drafted around it.

Envelope anatomy — every live envelope states all eight:

  1. Scope — the concrete work class it covers (a B3 row, or narrower).
  2. Magnitude — the largest single action (value, size, criticality tier).
  3. Velocity / rate — the volume budget: how many actions per period before the machine pauses itself.
  4. Reversibility window — how long an action can be undone at no or low cost, and by whom.
  5. Conditional no-gos — the concrete cases that always drop to human approval or never (the B3 row 12 list, asset criticality, first occurrence of a new pattern).
  6. Owner — the accountable person by name, with the power to halt.
  7. Expiry / re-review — the date the envelope stops being valid unless a person re-signs it.
  8. Edge behaviour — what the system does at the boundary: pause and hold, drop to approve-each, or quarantine the queue.

The trigger gate. An AI-initiated action passes the gate only if it is inside scope, under magnitude, under the volume budget, inside the reversibility window and not on the no-go list. Anything else is a recommendation, not an action.

The engineered stop. Someone named can pause the machine within the hour, the pause has been drilled in the last year, and the stop is designed to be cheaper than letting it run. Today: 36% of organisations have a named person who could pause within the hour, 22% have drilled it in the last year, 57% don't know ([W4] §3). On the W4 storm night the room let a queue of 14 run (47%), paused at 57 (71%) and split three ways at 428 — the pause point exists in people's heads and nowhere in their systems.

The volume budget — the community's draft number. Asked where auto-pause should kick in, in machine-created work orders per planner per week, the W4 room's median was 7.5 (n=30) — and 30% answered 0, never auto-create. Quote the median and the zero-share together; the mean (37.3) is dragged by a tail that ran to 300. The analogy is EEMUA 191's alarm budget (~1 alarm per operator per 10 minutes at steady state): a rate a person can actually oversee.

  • Volume budget (draft default): 7.5 machine-created work orders per planner per week
  • Review-period question: How often should a live envelope be re-reviewed and re-signed by its named owner?

The reversibility premium. The same pothole inspection drew 59% "allow auto-trigger" when cancellable for two hours and 21% when the cost committed on creation — 38 percentage points for one reversibility change (n=29/33). Reversibility moves the community more than work type does (five reversible classes spanning water, power, roads and health varied only 42–53%). Design the envelope to buy reversibility — cancellation windows, held-not-committed spend, quarantine — and the permitted level rises with it.

Minimum viable version. One work class, one volume cap, one cancellation window, one named owner, one tested stop.

Provenance: Synthesis F1–F3, F5 · [W4] §3–7 · W4 research pack dossier 05 (EEMUA 191) · [W5B] §1 axis (consequence × reversibility × evidence).


Confidentiality: every figure above is a consented, aggregated series result; no volunteer's cycle-1 text is reproduced. Cycle-1 inputs that shaped this draft are credited in the cycle-1 disposition log.

Cycle 2 — Cover Brief

Cycle 2 — Cover Brief

Stream 1 · Asset Zero. Drafted 17 Aug 2026 for the cycle-2 packet, live in Fulcrum from the early hours of Mon 17 Aug. AI-drafted from the series evidence and the W5 synthesis brief; Kai approves. Note on numbering: B3, B4 and B5 in this packet are framework sections (Part B — The boundary), not principle numbers.


What cycle 2 closes

The instruments — how the boundary is actually applied: the Decision Rights Matrix (B3), the boundary test (B4) and envelopes, triggers and the stop (B5). Cycle 1 settled what the boundary is (the one finding, the seven principles, the ladder); cycle 2 settles the three tools an organisation uses to place it. Under review: B3-B4-B5-Draft-v0.1.md (the drafted matrix rows, the scenario and nine questions, the envelope defaults) — about five pages.

Carried in from cycle 1 and applied throughout: Structure B (records and accountability as separate principles), the term "accountable person" (accountability and capability; definition still owed — see the disposition log), and the drafting guardrail agreed on the 14 Aug call — borrow vocabulary that maps to current practice, do not prescribe internal organisational structure, keep definitions concrete.

Dates (all AEST)

WhenWhatWhoBudget
Mon 17 Aug (early)Cycle-2 asks open at /members/tasks — B3, B4, B5Working Group
Mon 17 – Thu 20 Aug, 5:00 pmMicro-asks open (the exact close shows on each task)Working Group≤40 min total
Thu 20 Aug, 12:00–13:00 (to be confirmed by the chair)The call — contested rows and questions only; the asks stay open until 5 pm for anything the call surfacesWG + chair1 h
Fri 21 AugSections close; disposition log + revised draft publishedChair

The reviewer stream is parked for cycle 2 by the chair's decision (reviewer packet 1's findings are being folded into the framework; a second case is not issued this cycle).

The three asks, in one line each

  1. B3 — judge each drafted matrix row: agree with the drafted AI level (0–5), or propose the level you would set — one line why, on every row.
  2. B4 — apply the boundary test to the packet scenario: nine questions, yes / no / depends (and on what).
  3. B5 — set the numbers: the volume budget before auto-pause, how often a live envelope is re-reviewed, and what should always stop the machine.

How the decision gets made (unchanged)

The chair closes each section on rough consensus"can anyone not live with this?" — an objection blocks only until it has been addressed in writing. Every answer appears in the comment-disposition log with a written response before the section closes; unresolved disagreement is recorded as dissent and ships with the framework. Nothing is official until a named human signs it off. Written input carries identical weight to anything said on the call.

What changed in the instruments after cycle 1

Cycle 1's biggest data problem was unreasoned "keep" votes — 21 of 39 keeps came from three people who wrote 370 characters between them. So in cycle 2 the one-line reason is asked on every option, including "agree". A short line is fine; an empty one is not.


Provenance: Cycle-Plan.md §2–3 · HANDOFF-Cycle-1-to-Cycle-2-2026-08-16.md §5 · 14 Aug call minutes §9, A13–A15. AI-drafted; Kai approves before send.

Cycle 2 — Working Group Micro-Asks (spec)

Cycle 2 — Working Group Micro-Asks (spec)

Stream 1 · Asset Zero. Three structured asks, ≤40 minutes total, open in Fulcrum at /members/tasks from the early hours of Mon 17 Aug until Thu 20 Aug 5:00 pm AEST (the live close shows on each task). Answers are editable until close; each edit is a new revision and the latest submitted answer is the response of record. Reasons are asked on every option this cycle (the cycle-1 lesson). AI-drafted; Kai approves.


B3 — Decision Rights Matrix review (~15 min)

Per drafted row: agree with the AI level, or propose the 0–5 you'd set — one line why, on every row.

For each of the twelve rows in B3-B4-B5-Draft-v0.1.md §B3:

  • Agree with the drafted level — one line why (required)
  • Propose a different level — pick 0–5, one line why (required)

Design note: the reason is required on "agree" as well as "propose". Cycle 1 showed that an unreasoned keep cannot be distinguished from a rubber stamp; a five-word reason is enough. Rows are shown with the drafted level and the "humans must retain" cell; the evidence column is in the packet.

B4 — Apply the boundary test (~10 min)

Apply the nine questions to the packet scenario.

Read the scenario in B3-B4-B5-Draft-v0.1.md §B4, then for each of the nine questions:

  • Yes · No · Depends — and if it depends, one line on what (required for "depends")

Design note: the answers test the test, not the scenario. Where the group finds a question unanswerable or redundant, that is the finding — say so in the line.

B5 — Envelopes, trigger and stop (~10 min)

The numbers that keep autonomous action inside human-set bounds — judge the drafted defaults.

  • Volume budget — machine-created work orders per planner per week before the machine pauses itself (drafted default 7.5; the W4 room's median)
  • Review period — how many days a live envelope may run before its named owner must re-review and re-sign it
  • What should always stop the machine — the conditions that drop any AI action to human approval or to never, in your words (required)
  • Note — anything else about the envelope anatomy (optional)

Design note: numbers are displayed distribution-safe (medians, minimum n) — no individual number is published.

Handling rules (Stream 1 internal)

  • Nothing volunteer-written is edited. Every answer receives a written response in the disposition log before the section closes.
  • The reviewer stream is parked this cycle; reviewer-only volunteers see no cycle-2 asks unless they RSVP'd "both".
  • Numbers from B5 publish only as medians with n ≥ the site's MIN_N.

Cycle 1 packet — what volunteers are working with

Published in full, per the build-in-public rule; the calendar above carries the dates. The disposition log publishes with the closed sections.

Cycle 1 Draft — The One Finding and the Seven Principles (v0.2)

Cycle 1 Draft — The One Finding and the Seven Principles (v0.2)

The Human–AI Boundary Framework for Asset Management · sections A2 + B2. AI-drafted 2 Aug 2026 from the community's webinar data; every figure verified against the results files. Status: DRAFT — this is the text you are asked to challenge. Nothing here is official until the Working Group closes it and a named human signs it off.

Evidence citations: [W1]–[W4] are the four results files (Webinar N .../results/Webinar-N-Results-AI.md); [SES] is the Series Evidence Synthesis; [W5B] the capstone brief. Paired figures (same participants before/after) are marked paired — they are the honest ones.


A2 — Why a boundary framework (the one finding)

Across four webinars and 2,796 recorded responses, one finding held on every decision domain — data, condition and risk, investment, operations:

AI reliably amplifies whatever you already have. It can make incomplete data look complete, an uncertain prediction look like a decision, an incomplete objective look optimal, and a risky action look routine. Which way it resolves is a governance choice, not a property of the technology. The human–AI boundary is the mechanism by which AI-enabled asset management succeeds or fails.

The community did not vote to keep AI out: across every abstract boundary vote, "no role" and "should not be used" took 0% [SES F1]. It voted to place AI — analyse and recommend widely; act only inside engineered, reversible, approval-gated envelopes; and never own purpose, values, risk acceptance or the stop.

And it measured the gap this framework exists to close. The same practitioners who demand approval gates and stop authority report that today only 23% of their organisations auto-create work orders from alerts, 36% have a named person who could pause machine-generated work within the hour, 22% have drilled that pause in the last year — and their median guess for how many AI-generated plans get audited is 0% [W4 §3 · W3 w3-p2-audit-guess · SES F7]. The framework's job is to make the boundary explicit, decision by decision, and walkable from where members actually are.

Provenance: [W5B] §1 · [SES] §0, F1, F7. For the WG: is this framing right, and is "the boundary is a governance choice" the sentence you would defend to your board?


B2 — The seven principles (Structure A, as drafted in the capstone brief)

P1 — AI should improve asset-management decisions, not just produce analytics

Draft statement. The test of any AI use in asset management is whether an AM decision gets better — faster where speed matters, better-evidenced, more accountable — not whether a model produced an output.

What the community's data says. This is the series' framing device rather than a voted principle: every webinar asked "what decision changes?" before "what can AI do?". The strongest indirect evidence is negative — the community's most-seen failure pattern is confident output with no decision made better by it: 75% have seen bad data → confident output [W1 q8-patterns, n=71].

In practice. Every AI use case is registered against the decision it changes; "insight" tools with no decision owner don't pass screening.

For the WG. Is this a principle — or the framework's test of purpose, stated once in the preamble? (Structure B moves it there; see the structure question below.)

Provenance: [W5B] §3.1 · [SES] F7.


P2 — AI amplifies data; it does not fix it

Draft statement. Data quality, provenance and the Verified / Inferred / Unknown distinction are governance preconditions for trusting any AI output. No AI-authored value enters an authoritative register without labelling, a confidence score and human sign-off.

What the community's data says. Rated 3.86/5 (n=71; 70% rated 4–5) as W1's principle [W1 q9-first-principle]. 75% have seen bad-data→confident-output; 45% the authority illusion; 28% silent fabrication; 25% data laundering [W1 q8-patterns]. Asked which attribute must never be inferred without validation, the room's top answers were all of them and criticality [W1 q7-attribute]. Top-upvoted worry of the whole series: "generating bad data and results" (▲29) [W1 q2-hope-worry].

In practice. Register fields carry provenance labels; AI may sometimes act on non-authoritative fields within bounds (max level 4), but authoritative records cap at Recommend (level 3) with named sign-off [W5B §4 rows 1–2].

For the WG. Is Verified/Inferred/Unknown labelling practicable in your register today — and if not, what's the minimum viable version?

Provenance: [W5B] §3.2 · [W1] · [SES] F7–F8.


P3 — An AI prediction is not a risk decision

Draft statement. AI can estimate condition, failure probability or risk. Before a prediction triggers action, a competent human validates calibration, data quality and drift (is the asset still in-distribution?), and a named owner accepts and records the residual risk. Accountability never transfers to the model.

What the community's data says. Rated 3.76/5 (n=46; 65% rated 4–5) [W2 principle-test]. On acting on a prediction the room capped AI at analyse-or-recommend (67%, n=42) with autonomy at 5% [W2 boundary-vote-green-score]. Asked who owns the risk when AI flags, the room's answers converge on a named human — asset owner, asset manager, "whoever normally owns the risk, regardless of AI involvement" [W2 discussion]. Note honestly: accept & record residual risk was mid-to-low in every cold retained-rights vote (47% → 33% → 22%, W1→W3) and re-priced only after the room watched consequences (+12pp paired, W4) [SES F8].

In practice. Prediction-triggered work carries a risk-acceptance record: who accepted, on what evidence, with what reconsideration date.

For the WG. Does the "named owner" role exist in your organisation today — and at what level should it sit for, say, a chiller failure prediction?

Provenance: [W5B] §3.3 · [W2] · [SES] F7–F8.


P4 — AI may optimise to a target; humans decide whether the target is right

Draft statement. Optimisation exposes trade-offs; humans own the value framework and the choice. Only a named human sets and re-signs the objective's weights; whoever is affected may see the frontier (near-optimal alternatives and what each trades away); and every deferral an optimiser proposes is logged as a named, dated risk acceptance.

What the community's data says. This principle bundles W3's three rated candidates: the signed objective 3.73 (n=33), the frontier right 3.39 (n=36 — the series' weakest), the deferral ledger 4.07 (n=30 — W3's best) [W3 §2]. Behaviour was sharper than the ratings: 85% would not have signed the vendor-default weights nobody in the room chose [W3 w3-s1-r5-sign-default]; 64% rejected the optimiser's silent trade [W3 w3-s2-r5-accept-trade]; and in only 39% of organisations could a named person show you today the value weights behind the capital plan [W3 w3-s1-r6-weights-written]. What the room valued when it held the pen: cost 26 · risk 25 · service 24 · equity 15 · carbon 10 (medians, n=45) [W3 §4].

In practice. The value model is an owned, signed, re-signed artefact — like accounts. Optimiser outputs ship with the frontier view and a deferral ledger.

For the WG. The frontier right rated weakest (3.39). Keep it inside P4, promote it, demote it to a practice note — or is it simply ahead of the tooling?

Provenance: [W5B] §3.4 · [W3] · [SES] F6.


P5 — Automation depends on consequence, reversibility and monitoring

Draft statement. AI may recommend widely but should trigger narrowly: inside a defined envelope (scope, magnitude caps, rate limits, reversibility window, owner, expiry), behind an approval gate by default, with an engineered and practised stop. High-consequence or irreversible actions stay behind a human gate; some stay human, full stop.

What the community's data says. W4's paired experiment: after watching a bounded agent act for ~40 minutes, the room moved +0.44 rungs (paired, n=25) — to act with approval (26%→60%), not autonomy (13%→9%) [W4 §1]. The reversibility premium: 38pp — the same pothole work order drew 59% "allow auto" when cancellable for two hours vs 21% when cost commits on creation [W4 §6]. The room drafted its first volume budget: median 7.5 machine-created work orders per planner per week before auto-pause; 30% said never [W4 §5]. Trigger gate rated 4.15, engineered stop 4.06 [W4 §9]. Stop authority was the only retained right nobody dropped after the storm (48%→64% paired) [W4 §1e]. And "never" votes appeared only when the case got concrete: 19% never on live 66kV isolation vs 0% never in every abstract vote [W4 §6 · SES F3].

In practice. Every acting AI has an envelope document and a named owner; the stop is a tested control, drilled like an emergency procedure, not a menu item.

For the WG. Does the envelope anatomy belong inside the principle, or in the framework's B5 section with the principle kept short?

Provenance: [W5B] §3.5 · [W4] · [SES] F1–F3, F5.


P6 — Human oversight must be meaningful and verifiable — and accountability cannot be delegated

Draft statement. "A human reviews it" is a claim, not a control. Oversight counts only when it is verifiable: interrogable outputs, visible limitations, effective challenge by people with the capability, authority and time to say no — and a named accountable owner with the power to halt. "The model said so" is no defence.

What the community's data says. The community's confidence in declared oversight is low: median guess for the share of AI-generated plans audited today — 0% [W3 w3-p2-audit-guess]. 14% have already seen rubber-stamp oversight; 45% the authority illusion [W1 q8-patterns]. Asked which safeguard their organisation would quietly skip first, 37% said none — and among the clauses actually named, the close-out feedback loop (30%) and the practised stop drill (22%) topped the list, while 0% believed the re-signing schedule would be skipped [W4 §9]. The room's own behaviour under pressure: approval and stop rights rose (paired +12pp, +16pp), abstract stewardship fell [W4 §1e · SES F2, F8].

In practice. For high-consequence decisions: challenge is resourced and recorded; oversight quality is audited (spot-checks of approvals), not assumed; every AI use case has one named accountable owner who can halt it.

For the WG. What makes oversight verifiable in your world — records (the ledger), independent challenge, or drills? Which would you mandate first?

Provenance: [W5B] §3.6 · [SES] F2, F4, F7.


P7 — AI should augment professional judgement, not deskill it

Draft statement. Design AI use so the expertise the boundary depends on keeps being exercised: rotation through manual practice, visible reasoning rather than bare answers, and monitoring for out-of-the-loop erosion — because every earlier principle assumes a human who can still meaningfully validate, challenge and stop.

What the community's data says. This is the series' least-quantified principle — flag that honestly. The signal is in the open text (verbatim): "Loss of abilities and skills to solve problems", "Cognitive offloading", "Loss of ability to evaluation AI responses, true /false", "Will management see this as a way to remove our skill?" [W1 q2, q10], and deskilling stands as an open question in the capstone register [W5B §9]. No boundary vote tested it.

In practice. Capability impact is assessed at screening; augmentation patterns (AI drafts, human decides; AI explains, human verifies) are preferred over replacement patterns for judgement-bearing tasks.

For the WG. Least-evidenced principle in the set: keep it as a principle, or convert it to capability practice-notes under C3 — and if it stays, what evidence should the final pass (or a v1.x cycle) gather?

Provenance: [W5B] §3.7 · [W1] · [SES] F8.


The structure question (MA1 asks you to pick)

The seven above are Structure A — the capstone brief's set, one principle per webinar plus three cross-cutting. The series' strongest single result sits awkwardly inside it: the ledgers won their rooms (action ledger 4.27 = the series' best; deferral ledger 4.07 = W3's best), yet "keep a record" appears only as clauses inside P4/P6. Structure B makes it load-bearing:

Structure A (capstone draft)Structure B (ledger elevated)
A1 · Decisions, not analyticsbecomes the preamble: the framework's test of purpose
A2 · Data: amplifies, doesn't fix (V/I/U)B1 · Data: amplifies, doesn't fix (V/I/U)
A3 · Prediction ≠ risk decisionB2 · Prediction ≠ risk decision
A4 · Optimise to a target; humans own the targetB3 · Values: humans own the target and the trade-offs
A5 · Automation: consequence × reversibility × monitoringB4 · Action: envelopes, approval gates and the engineered stop
A6 · Oversight meaningful & verifiable + accountability non-delegableB5 · No action without a record — decision records, action & deferral ledgers, re-signing (records are what make oversight verifiable)
A7 · Augment, don't deskillB6 · Accountability: named, non-delegable, with the power to halt
B7 · Capability: augment, don't deskill

Both structures hold seven principles: B moves A1 to the preamble and splits A6 into records (B5) and accountability (B6). The trade: A maps one-to-one onto the webinar journey; B makes the community's strongest-rated mechanism load-bearing and gives accountability its own line.

Evidence for elevating records: P4-C 4.27 (82% rated 4–5) · P3-C 4.07 (83%) · 0% believed the re-signing schedule would be skipped · the community consistently endorsed recording mechanisms above access rights and process promises [SES F4]. Evidence for caution: Structure A maps one-to-one onto the webinars (traceability story is cleaner), and the brief's set is already public in draft.

Provenance: [SES] F4 · [W5B] §3 · Framework-v0.1-Skeleton §3.


AI-use statement: this draft was AI-synthesised from the community's consented webinar responses and the capstone brief; figures were verified against the results files in two fact-check passes. The Working Group judges; a named human signs. Comment disposition: every comment on this draft receives a written answer in the cycle-1 log.

Cycle 1 — Cover Brief

Cycle 1 — Cover Brief

Stream 1 · Asset Zero. Drafted 2 Aug 2026; re-dated 8 Aug for the Fulcrum-only capture decision (kickoff moved to Tue 11). This is page 1 of the cycle-1 packet, live in Fulcrum from Tue 11 Aug.


What cycle 1 closes

The framework's spine: the One Finding (A2), the Seven Principles (B2), and the Ladder & Human Roles (B1) — plus one structural decision: which of two principle-set structures the framework uses. Everything later builds on what you settle here. Under review: B2-A2-Draft-v0.2.md (7 principle cards + the structure question, ~6–7 pages) plus the B1 ladder question (the series' 0–5 ladder as carried in the framework skeleton; the one-page B1 ladder/roles draft follows).

Dates (all AEST)

WhenWhatWhoBudget
Tue 11 Aug, 12:00–12:30Kickoff call (recorded) · cycle-1 asks open in FulcrumAll30 min
Tue 11 – Thu 13 Aug (noon)Micro-asks MA1–MA3 open at /members/tasksWorking Group≤40 min total
Fri 14 Aug, 12:00–13:00The call — contested points onlyWG + chair1 h
Tue 11 – Fri 14 AugReviewer packet ("The Confident Wrong Answer")Reviewers≤1 h
Mon 17 Aug (AM)Sections close; disposition log + revised draft publishedChair

How the decision gets made (once, so it never surprises anyone)

The chair closes the section on rough consensus: not a vote count, but the test "can anyone not live with this?" — an objection blocks only until it has been addressed (accepted, or answered with reasons in writing). Every comment you make — form, email or call — appears in the comment-disposition log with a written response before the section closes. Unresolved disagreement is recorded as dissent and ships with the framework; it is data, not failure. Nothing is official until a named human signs it off.

The call agenda (60 min)

  1. 5' — Welcome; the method above, said out loud once.
  2. 40' — Contested points only, from your MA1/MA2 answers (expect: the structure question, the frontier right P4, and whatever MA2 breaks hardest).
  3. 10' — Consensus test, principle by principle (B1 ladder included).
  4. 5' — Close: what ships, what carries dissent, cycle-2 preview (the Decision Rights Matrix, boundary test and envelopes — Thu 20).

Can't make it? The call is recorded, a 15-minute catch-up pack follows, and your written input carries identical weight — decisions close on the log, not in the room.

Credit

Working Group members are named co-developers on the framework's cover; Reviewers are named on the contributors page; all webinar participants are acknowledged collectively. Details: the contributor one-pager sent with the kickoff email.


Provenance: Cycle-Plan.md §3–5 · R04 §E. AI-drafted; Kai approves before send.

Comment-Disposition Log — Template (all cycles)

Comment-Disposition Log — Template (all cycles)

Stream 1 · Asset Zero. Drafted 2 Aug 2026. One log per cycle, published with the closed section. The promise it operationalises: every input gets a written answer. (ISO comment-disposition discipline, R04 §A/§C.)


Columns

ColumnRule
#C1-001 … sequential within the cycle
ContributorName, or "anonymous" on request — never blank
SourceMA1 · MA2 · MA3 · call · reviewer-Q1/Q2/Q3 · email
Section / principleWhat the comment targets (e.g. B2-P4, structure question)
CommentVerbatim — never edited, only truncated with "[…]" if very long (full text archived)
DispositionAccept (text changed as asked) · Partial (changed, differently — say how) · Decline (with reasons) · Defer (moved to Open Questions Register or a later cycle, with ID)
ResponseThe written answer — 1–4 sentences, specific, citing evidence where it decides the matter
ResolverAI-drafted / chair-adjudicated — final column always carries the chair's initials
DateClosed date

Rules

  1. AI drafts dispositions; the chair adjudicates every row — no row closes on AI authority.
  2. Identical comments are grouped (one row, all contributors listed) — but each contributor is named.
  3. Decline requires reasons that cite evidence or a principle, never "out of scope" alone.
  4. Defer rows must land somewhere visible: an Open Questions Register ID or a named future cycle.
  5. The log is published with the closed section and archived in the repo (cycles/cycle-N/disposition-log.md); it feeds the framework's D4 provenance appendix.
  6. A section may close only when every row has a disposition — this is the chair's checklist, not a formality.

Example row

#ContributorSourceSectionCommentDispositionResponseResolverDate
C1-007J. Example (Utility X)MA2B2-P5"Rate limits assume a planner sees the queue daily; our depot checks weekly — 7.5/week is meaningless there."PartialBudget reframed as per-planner-per-review-period with 7.5/wk as the weekly default; example added. Weekly-review case logged to register as OQ-refinement.AI-draft / KD14 Aug

Provenance: R04 §A (ISO 37000: every ballot comment answered individually), §C, §E. Template is cycle-agnostic; copy per cycle.

Cycle 1 — Working Group Micro-Asks (spec)

Cycle 1 — Working Group Micro-Asks (spec)

Stream 1 · Asset Zero. Drafted 2 Aug 2026; re-dated 8 Aug for the Fulcrum-only capture decision (kickoff Tue 11). Three asks, ≤40 minutes total, open Tue 11 Aug – close Thu 13 Aug noon (AEST). Answered in Fulcrum at /members/tasks (sign-in required; answers autosave and stay editable until close).


MA1 — Judge each principle, and pick the structure (~15 min)

For each of the seven principles in the draft (P1–P7), one required choice and one optional line:

P[n]: [short title]Keep as drafted ☐ Sharpen (what would you change?) ☐ Reject (why?) One line — the reason for your call: ______

Then one structural choice:

The structure question (see the draft's final section): ☐ Structure A — the capstone set, one principle per webinar ☐ Structure B — records elevated ("no action without a record"), accountability separated ☐ Either, with this condition: ______

Design notes: radio + one free-text line per principle — nothing open-ended beyond that (Delphi fatigue rule, R04 §B). Progress bar. Autosave. Est. completion ~15 min. The B1 ladder/roles page carries one added radio: does the 0–5 ladder read right for your sector?

MA2 — Break one principle (~20 min)

Pick the ONE principle you know best from your own work. Describe a real situation from your organisation (anonymised — no org names, no identifying details) where the principle as drafted would fail, mislead, or be quietly ignored. Three prompts:

  1. The situation (3–5 sentences).
  2. Where exactly the principle breaks (one sentence).
  3. What wording or condition would fix it (one sentence — or "it can't be fixed because…").

Design notes: this is the cycle's highest-value ask — it generates the failure-mode and practice-note material for the cycle-3 final pass (C2/C3). One entry required, more welcome. Anonymised entries feed the case library with contributor consent per the one-pager.

MA3 — The missing principle (~5 min)

If you could add an eighth principle, what would it say? One sentence in principle form ("AI may… / humans must…"). If nothing is missing, say so — "the set is complete" is a valid and useful answer.


Handling rules (Stream 1 internal)

  • Every response lands in the cycle-1 comment-disposition log within 24 h, verbatim, tagged MA1/MA2/MA3 + contributor (or "anon on request").
  • MA1 tallies + MA2 themes become the call's contested-points list (built Thu 13 Aug afternoon, sent to WG before Friday's call).
  • Non-responders get exactly one nudge (Thu 13 Aug afternoon, from Kai, personal, two lines). Watch threshold: <60% by Thu noon (Cycle-Plan v2 §6).
  • Nothing volunteer-written is edited — dispositions respond to the original text.

Provenance: R04 §E workflow · Cycle-1-Packet-Outline §3 · B2-A2-Draft-v0.2 (the object under review). AI-drafted; Kai approves before send.

Cycle 1 — Reviewer Packet: "The Confident Wrong Answer"

Cycle 1 — Reviewer Packet: "The Confident Wrong Answer"

Stream 1 · Asset Zero. Drafted 2 Aug 2026; re-dated 8 Aug for the Fulcrum-only capture decision. Open Tue 11 Aug · due Fri 14 Aug (AEST) · ≤60 minutes. One case, three questions.

Reviewers: you're testing whether the draft principles survive contact with a realistic case. Your answers go, verbatim, into the comment-disposition log — every one gets a written response before the section closes.


The case (composite — fictional organisation, built from patterns the community reported)

Regional water utility, ~40,000 assets. Eighteen months ago it licensed an AI asset-intelligence platform. Three things happened, in order:

1 — The register got "complete". The platform back-filled missing attributes across the register — install dates inferred from mains-laying programs, materials inferred from era and suburb, condition scores extrapolated from the 8% of the network CCTV'd in the last decade. The enriched register looked authoritative: every field populated, professional confidence scores in a tooltip nobody hovers. Within a quarter, the inferred values were being exported into renewal models and board papers with no provenance flags. (The community has met this: 75% have seen bad data → confident output; 28% silent fabrication; 25% data laundering — [W1 q8-patterns], n=71.)

2 — Criticality moved. The platform re-scored asset criticality network-wide. A trunk main serving the hospital precinct — never inspected, no failure history because it had always been quietly managed — scored "low criticality, defer". The score was arithmetically defensible from the data it had. The data was the problem. (Criticality was the community's #1 named never-infer-without-validation attribute — [W1 q7-attribute].)

3 — The answer sounded like the senior engineer. Planners began querying the platform's assistant directly. Asked about the trunk main, it blended the inferred condition, the superseded 2019 renewal strategy and the current one into a fluent recommendation to defer — delivered with the same confident tone it used when it was right. The one engineer who remembered why that main mattered retired in March. (The authority illusion: 45% have seen it — [W1 q8-patterns].)

Nobody has acted yet. The deferral is in this year's draft plan, four weeks from sign-off. You have been asked to review the plan.


Your three questions

Q1 — Catch it. Which of the seven draft principles (P1–P7), applied as written, would have caught or prevented each of the three failures — and at which step does the draft wording actually fail to bite? Be specific: "P2 catches failure 1 at the register boundary, but says nothing about exports" is the kind of answer that changes the text.

Q2 — Reality-test it. For the principle you rate most load-bearing here: does your organisation run the control that principle assumes (provenance labels, a named criticality owner, challenge of fluent AI answers)? If not, what is the smallest version that would have caught this case?

Q3 — One sentence for the framework. Write the one sentence you would put in the framework's practice notes so that a planner four weeks from sign-off, in this exact position, does the right thing.

Format: free text, any length (a paragraph per question is plenty). Answered in Fulcrum at /members/tasks (autosaves; editable until the packet closes). Anonymised excerpts may be used in the framework's case library per the contributor one-pager — flag anything you'd rather keep out.


Provenance: composite case AI-drafted from [W1] q8-patterns, q7-attribute, q2-hope-worry and the consolidated failure-modes library ([W5B] §6: bad-data/confident-output, silent fabrication, data laundering, the unofficial authority, biased risk scores). No real organisation is depicted; the failure patterns are the ones the community itself reported. Kai approves before send.