Series Evidence Synthesis — What Four Webinars Actually Said
Stream 1 · Project Output · Asset Zero. Drafted 2 Aug 2026 (AI-drafted, fact-checked against the results files; provisional until a named human signs it off). This is the spine document: every framework section drafts from here, and every claim here traces to a results file. Nothing in this document is an official output until a named human signs it off.
0. Sources and how to read this document
| Key | File (repo path) | Session | Date | Scale |
|---|---|---|---|---|
| [W1] | Webinar 1 - Knowing the Asset/results/Webinar-1-Results-AI.md (+.json, xlsx) | Knowing the Asset | 25 Jun 2026 | 663 responses · largest single-question n=98 (Mentimeter import) |
| [W2] | Webinar 2 - Condition, Performance and Risk/results/Webinar-2-Results-AI.md (+.json) | Condition, Performance & Risk | 1 Jul 2026 | 524 responses (first live Fulcrum run) |
| [W3] | Webinar 3 - Planning, Prioritisation and Investment/results/Webinar-3-Results-AI.md (+.json) | Planning, Prioritisation & Investment | 15 Jul 2026 | 63 participants · 625 responses (+28 sector answers outside the cue script) |
| [W4] | Webinar 4 - Work, Operations and Intervention/results/Webinar-4-Results-AI.md (+.json) | Work, Operations & Intervention | 30 Jul 2026 | 60 participants · 956 responses |
| [W5B] | Webinar 5 - Synthesis and Recommendations/Webinar-5-Synthesis-Brief.md | Capstone brief (draft v1.0 outputs) | — | Synthesis target, not evidence |
Series total: 2,796 recorded responses (663 + 524 + 625 + 956 in the live bundles, plus 28 W3 sector answers logged outside the cue script). Use "almost 2,800 responses" in comms; do not round to 3,000 in anything citable.
Reading rules (they are printed in the files; they bind here too):
- Every figure carries its own n. Rooms shrink across an hour; compare within a block, not across blocks.
- Where a paired figure exists (same participants matched on anonymised token), it is the honest figure and the one to quote: W4 boundary ladder (+0.44, n=25), W4 retained-rights re-price (n=25), W4 comfort (+0.18, n=17), W3 trust (+0.43, n=14).
- Multi-select percentages are % of that block's voters and sum past 100%.
- Percentages below use each file's own rounding. Ratios computed on raw counts.
- Cross-webinar comparisons are cross-room (different people, different sector mixes) except where an instrument was deliberately re-run verbatim ([W4] §7 re-runs [W2]
dr-poll-c_trig; [W4] C04/C21 re-runs the series boundary vote twice in one room).
1. The series at a glance
| W1 · Knowing the Asset | W2 · Condition/Perf/Risk | W3 · Planning/Investment | W4 · Work/Operations | |
|---|---|---|---|---|
| Core question | Can AI help us know our assets, or amplify poor data? | When AI predicts, what must humans validate before acting? | Should AI help decide where the money goes — and who owns that call? | When should AI move from recommending work to triggering work? |
| Boundary-vote ceiling (analyse+recommend) | 68% (n=78) | 67% (n=42) | 72% (n=18) | cold 52% → warm 32% — the room moved to act with approval (n=39/35) |
| "Act with approval" share | 15% | 19% | 17% | cold 26% → warm 60% |
| "Act autonomously within limits" | 8% | 5% | 0% | cold 13% → warm 9% |
| "No role" / "Should not be used" | — / 0% (no "no role" option) | 0% / 0% | 0% / 0% | 0% / 0% (both legs) |
| Principle rating(s) /5 | 3.86 (n=71) | 3.76 (n=46) | A 3.73 · B 3.39 · C 4.07 | A 4.15 · B 4.06 · C 4.27 |
Sources: [W1] q5-ai-role, q9-first-principle · [W2] boundary-vote-green-score, principle-test · [W3] §1a, §2 · [W4] §1a–1b, §9.
Room composition where measured: W3 was 43% transport/roads, 36% consulting (n=28, [W3] §6); W1 skewed government/infrastructure/transport ([W1] q3-sector). W4 did not capture sector (its C06 slot went to the org-reality questions, [W4] §3). Treat sector mix as a confound on all cross-webinar reads.
2. Finding 1 — The ladder holds at Recommend, until the room watches bounded action work
For three webinars, on three different decision domains, the community capped AI at analyse-or-recommend by a stable supermajority: 68% on facts about the asset ([W1] q5-ai-role, n=78), 67% on acting on a prediction ([W2] boundary-vote-green-score, n=42), 72% on money ([W3] §1a, n=18). Autonomy peaked at 13% (W4 cold, 5/39) and hit 0% on money. And in all four series boundary votes (W2, W3, W4 cold & warm), "no role" and "should not be used" took 0% every time (W1's poll, which offered only "should not be used", likewise 0%); across the eight decision-rights polls (W2×5, W4×3), "don't use AI" never drew more than one vote (≤3%). The community, in the abstract, never votes to ban AI (concrete "nevers" do exist — see Finding 3).
W4 ran the series' controlled experiment: the identical boundary vote, cold at C04 and warm at C21, after the room spent ~40 minutes watching a bounded agent act (the storm, the envelope, the 2 a.m. signature). Result:
| Cold C04 (n=39) | Warm C21 (n=35) | |
|---|---|---|
| Assist only | 10% | 0% |
| Analyse | 21% | 6% |
| Recommend | 31% | 26% |
| Act with human approval | 26% | 60% |
| Act autonomously within limits | 13% | 9% |
Paired (n=25): mean ladder move +0.44 steps — 10 moved AI up, 5 down, 10 stayed ([W4] §1c). The warm room did not move toward autonomy (13%→9%); it moved from analyse/recommend to act with approval.
Spine claim: the community's default is Recommend; demonstrated, bounded, stoppable, ledgered action moves it one step — to approval-gated action, not autonomy. The ladder's load-bearing rung is Act with approval, and it is earned by showing the controls, not by asserting them.
Corroborating trend on the one verbatim re-run instrument — the intervention trigger ([W2] dr-poll-c_trig n=41 → [W4] §7 n=33, different rooms, 4 weeks apart): Recommend 71%→33%, Act-with-sign-off 5%→24%, autonomous 0%→0%. Same shape: the middle of the ladder is migrating one rung, autonomy is not moving.
3. Finding 2 — Approval over interruption: the gap that widened for three webinars, and W4's re-price of the stop
The community consistently prices approving before far above interrupting after. Define the gap as Approve the decision ÷ Hold stop authority (W1 offered no stop option; its nearest interrupt right, Override, is used for W1 only):
| Approve | Stop authority | Gap | |
|---|---|---|---|
| W1 (n=79; override as proxy) | 66% | 61% (override) | ~1.1× |
| W2 (n=42) | 48% (20) | 29% (12) | 1.7× |
| W3 (n=18) | 44% (8) | 17% (3) | 2.7× (3.3× vs own the value judgement, the W3 top right at 56%) |
| W4 cold (n=39) | 67% | 49% | 1.4× |
| W4 warm, paired (n=25) | 80% | 64% | 1.25× |
Sources: [W1] q6-humans-retain · [W2] boundary-vote-green-score (humans retain) · [W3] §1b · [W4] §1d–1e. W3's n=18 is small; treat its ratio as directional.
Stop authority and override ranked at or near the bottom of every humans-must-retain vote from W2 to W3 (W2: 12–13 of ~42 voters; W3: 17%). This was the tension W4 was built to test — the session showed what the stop costs (the storm; EEMUA-style flood; the quarantine call), then re-ran the vote.
The re-price ([W4] §1e, paired n=25): stop authority 48% → 64% — 4 added it, 0 dropped it. It is the only retained right that nobody dropped after watching the machine act. Approval also rose (68%→80%); risk acceptance rose (40%→52%); interpret-context doubled (20%→40%). What fell: govern-the-knowledge (60%→48%) and own-the-value-judgement (48%→44%) — under pressure the room traded abstract stewardship rights for concrete interrupt-and-accountability rights.
Spine claim: approval is the community's currency, but the stop was systematically under-priced until the room saw a storm. The framework must price the stop explicitly (P4-B) — and the evidence says demonstration, not exhortation, is what re-prices it.
Reality check (Finding 7 has the full set): only 36% have a named person who could pause machine-generated work within the hour today; the pause has been drilled in the last year in 22% of organisations, and 57% don't know ([W4] §3).
4. Finding 3 — Reversibility is worth 38 points; "never" only appears when the case is concrete
W4's envelope exercise put six concrete work classes on an allow / approval-gate / never ballot ([W4] §6):
| Work class | Allow auto-trigger | Approval gate | Never |
|---|---|---|---|
| E1 pothole make-safe inspection (n=33) | 45% | 55% | 0% |
| E2 substation thermal re-inspection (n=31) | 42% | 58% | 0% |
| E3 sewer CCTV follow-up (n=32) | 53% | 47% | 0% |
| E4 hospital chiller filter replacement (n=36) | 44% | 56% | 0% |
| E5 spares reorder <$500 (n=36) | 47% | 50% | 3% |
| E6 live 66kV feeder isolation + permit to work (n=36) | 11% | 69% | 19% |
The reversibility flip ([W4] §6, the premium experiment): the same pothole inspection drew 59% allow when cancellable for two hours (n=29) vs 21% allow when cost commits on creation (n=33) — a 38-percentage-point reversibility premium, measured on this community for the first time. (Note the two flip variants had different n; the premium is a cross-block read within one session and one instrument family.)
Two structural observations:
- Abstract bans never win; concrete bans exist. Across every abstract boundary vote, "should not be used" took 0%. Put a live 66kV isolation in front of the same community and 19% say never, and allow-auto collapses to 11%. The framework's "never automate" list must be written in concrete work classes, not categories of principle.
- Reversibility moves votes more than work type does. The six work classes span water, power, roads, health — allow-auto varies only 42–53% across five of them. One reversibility change moves the same class by 38pp. This validates the W5 brief's axis — consequence × reversibility × evidence ([W5B] §1) — with reversibility as the strongest measured lever.
Spine claim: the Decision Rights Matrix's "Act (bounded)" column is really a reversibility column. Envelope anatomy (scope, magnitude caps, rate limits, reversibility window, owner, expiry) is what the community is actually pricing.
5. Finding 4 — The community trusts records over promises: ledger shapes rate highest
Eight principles were rated live across the series (1 = reject, 5 = adopt):
| Principle (short) | Session | Mean | n | Rated 4–5 |
|---|---|---|---|---|
| Verified / Inferred / Unknown + human sign-off | W1 q9 | 3.86 | 71 | 70% |
| Prediction ≠ decision; named owner accepts residual risk | W2 principle-test | 3.76 | 46 | 65% |
| P3-A · The signed objective | W3 §2 | 3.73 | 33 | 73% |
| P3-B · The frontier right | W3 §2 | 3.39 | 36 | 47% |
| P3-C · The deferral ledger | W3 §2 | 4.07 | 30 | 83% |
| P4-A · The trigger gate | W4 §9 | 4.15 | 33 | 79% |
| P4-B · The engineered stop | W4 §9 | 4.06 | 35 | 71% |
| P4-C · The action ledger | W4 §9 | 4.27 | 33 | 82% |
The ledgers win their rooms: the highest-rated principle of the whole series is the action ledger (P4-C, 4.27 — every machine action logged with an owner's signature), and W3's highest was its twin, the deferral ledger (P3-C, 4.07 — every deferral logged as a named, dated risk acceptance); each beat every alternative rated alongside it. (P4-A, the trigger gate, sits between them at 4.15.) The weakest of the series, P3-B (3.39), is the one principle that grants a right to see rather than requiring a record to exist.
Consistent signals from W4's honesty poll ("which clause would your organisation quietly skip first?", n=27, [W4] §9): re-signing schedule 0% — nobody thinks the signature ritual would be skipped; what dies first is the close-out feedback loop (30%) and the practised stop drill (22%) (37% claim they'd keep all). The community believes in signatures and doubts follow-through — precisely the ledger-over-promise pattern.
Spine claim: wherever the framework must choose a mechanism, choose the recording mechanism: named, dated, signed, auditable. The community endorses accountability artefacts more strongly than access rights, oversight declarations, or process promises. (A consistent reading, offered as interpretation: W2's principle — strong content, no ledger clause — landed mid-pack at 3.76.)
6. Finding 5 — The community set its first number: 7.5 machine-created work orders per planner per week
Asked where auto-pause should kick in — machine-created work orders per planner per week — the room drafted the missing volume budget ([W4] §5, n=30):
- Median 7.5 · mean 37.3 (the mean is dragged by a heavy tail: answers ran 0→300)
- 30% answered 0 — never auto-create
- Full distribution is published in [W4] §5 (all 30 answers, sorted)
This is the community's first quantified boundary parameter — the seed of an "EEMUA-for-work-orders" style volume budget (the alarm-management analogy: EEMUA 191's steady-state budget of ~1 alarm per operator per 10 minutes; see the W4 research pack, dossier 05). Quote it as "the room's median was 7.5; a third said never" — the median and the zero-share together are the finding; the mean alone misleads.
Spine claim: the framework's envelope section can ship a worked example with a community-sourced default rather than a blank. Status: draft parameter for validation in cycles (OQ2 in the open-questions register), not a recommendation yet.
7. Finding 6 — Money is where values hide: weights, defaults and the signed objective
W3 put the value-framework question directly to the room:
- What this room valued (100 points across five criteria, n=45, [W3] §4): medians — cost 26 · risk 25 · service 24 · equity 15 · carbon 10, last. Cost/risk/service take ~three-quarters of the weight. (No is/ought baseline exists — the pre-survey never opened; flag when quoting.)
- 85% would not have signed the vendor-default weights nobody in the room had chosen (n=39, [W3]
w3-s1-r5-sign-default). - Where are your organisation's value weights written down? (n=36): a named person could show me them today 39% · a team could reconstruct them 25% · we don't use a weighted model 19% · nobody 11% · only the vendor 6% ([W3]
w3-s1-r6-weights-written). - Shown what an optimiser quietly defunded, 64% rejected the trade (n=42, [W3]
w3-s2-r5-accept-trade); asked who owns checking that today: we don't use an optimiser 37% · a named person 32% · a team could 17% · nobody 12% · only the vendor 2% (n=41). - Median guess for the share of AI-generated plans that get audited: 0% (mean 10%, n=35, [W3]
w3-p2-audit-guess). - Sim 3 (sign / send back / override on four AI-assembled plans): no plan reached a majority to sign; the strongest response to the "beautiful but gamed" plan was send-back/override 72% combined ([W3] §5). The room's instinct on machine-assembled plans is scrutiny, not signature.
Spine claim: P3-A (only a named human sets and re-signs the weights) is evidenced not by what the room endorsed (3.73) but by what it did: refusing default weights 85%, rejecting the optimiser's silent trade 64%, and reporting that in most organisations the weights and the check have no named owner. The matrix rows for prioritisation/valuation must carry the named-accountability column as their load-bearing cell.
8. Finding 7 — The is/ought gap: rooms demand gates their organisations don't run
Put the boundary votes (ought) beside the org-reality polls (is):
| The community demands (ought) | The community reports (is) |
|---|---|
| The approval gate drew 47–69% across the six envelope classes — the most-chosen option on five of six ([W4] §6) | 23% auto-create work orders from alerts today ([W4] §3, n=40) |
| Stop authority re-priced to 64% paired ([W4] §1e) | 36% have a named person who could pause within the hour; 22% No; 28% don't know ([W4] §3, n=36) |
| P4-B engineered stop rated 4.06 ([W4] §9) | Pause drilled within the last year: 22%; never 14%; don't know 57% ([W4] §3, n=37) |
| P3-A signed objective rated 3.73 ([W3] §2) | Weights written down and showable by a named person: 39% ([W3] w3-s1-r6) |
| Oversight must be verifiable ([W5B] P6) | Median guess of AI-plan audit rate: 0% ([W3] w3-p2-audit-guess) |
| In the storm, 71% paused by frame 2 ([W4] §4, n=31) | Who could actually pause this? Named-within-the-hour 42%, nobody 13%, don't know 26% ([W4] §4 F2b, n=31) |
Also W1's field-validation poll ([W1] q8-patterns, n=71, multi-select): 75% have seen "bad data → confident output", 45% the authority illusion, 28% silent fabrication, 25% data laundering, 17% rare-event blindness, 14% rubber-stamp oversight; only 11% "none yet". The failure-modes library is not hypothetical — the room has met it.
Trust starts low and moves slowly: trust in an AI-generated capital plan 2.64/7 before → 3.12 after, paired +0.43 (n=14) ([W3] §3); comfort with the overnight work order 2.86/7 → 3.11, paired +0.18 (n=17) ([W4] §2). One session of evidence and controls buys about a third of a point on a seven-point scale. Trust is earned in inches — which is the argument for the living framework's telemetry over a one-off PDF.
Spine claim: the framework's job is to close an is/ought gap the community itself measured. That is the "getting started → maturing" path's evidence base ([W5B] §7.1, §7.6): most member organisations are at the left edge (no auto-creation, no drilled stop, unwritten weights), while their practitioners already demand the right-edge controls. Both facts are in the data; the framework must serve both.
9. Finding 8 — The human core is stable across the lifecycle — and it sharpens under pressure
What humans must retain, across every instrument that asked (multi-select; % of that block's voters; instruments evolved, so read ranks within a column, not levels across columns):
| Retained right | W1 (n=79) | W2 (n=42) | W3 (n=18) | W4 cold→warm paired (n=25) |
|---|---|---|---|---|
| Approve the decision | 66% | 48% | 44% | 68% → 80% |
| Validate the evidence/data | 59% | 48% | 33% | 52% → 64% |
| Own the value judgement | — | 38% | 56% | 48% → 44% |
| Accept & record residual risk | 47% | 33% | 22% | 40% → 52% |
| Hold stop authority | — | 29% | 17% | 48% → 64% |
| Override | 61% | 31% | 17% | 56% → 60% |
| Interpret context | 51% | 38% | 17% | 20% → 40% |
| Govern the knowledge | 66% | 29% | 39% | 60% → 48% |
| Define purpose | 57% | — | — | — |
Sources: [W1] q6-humans-retain (7-option list: no stop-authority or own-value options; had define-purpose) · [W2] boundary-vote-green-score (8-option list stabilises here) · [W3] §1b · [W4] §1d–1e.
Stable structure: Approve is top-two in every room that voted it. Validate is always high. The domain flavours the top right — W3 (money) uniquely elevates own the value judgement to #1; W1 (data) elevates govern the knowledge. And W1's weakest right — accept risk, 47%, the lowest of its seven — stayed weak in W2–W3 (33%, 22%) until W4's action session re-priced it (+12pp paired). The pattern of Finding 2 generalises: exposure to consequence shifts the community from stewardship rights to interrupt-and-own rights (interpret-context +20pp, stop +16pp, approve +12pp, accept-risk +12pp; govern −12pp, own-value −4pp).
W1's open text corroborates the core: top-upvoted hopes/worries are bad data and results (▲29), lack of human oversight (▲24), data security (▲19), and one answer asks, in effect, who goes to jail for an issue controlled by AI ([W1] q2-hope-worry) — accountability by name, again. The questions the room left with centre on accuracy and trust ([W1] q10-leaving-question).
Spine claim: the framework's human-role set (define purpose · validate evidence · approve the decision · own the value judgement · accept & record residual risk · hold stop authority · override · interpret context · govern the knowledge · remain accountable by name) is fully evidenced across four rooms — with measured guidance on which rights need defending in the text (the ones rooms under-price cold: stop, risk acceptance, context) versus which defend themselves (approval, validation).
10. What the spine means for each framework section
| Finding | Feeds framework section (v0.1 skeleton) | Cycle |
|---|---|---|
| F1 ladder holds at Recommend; approval is the earned rung | B1 The ladder & human roles · B3 Decision Rights Matrix | 2 |
| F2 approval over interruption; the stop re-priced | B2 principles (P-Stop) · B5 trigger gate & stop · C3 verifiable oversight | 1, 3, 5 |
| F3 38pp reversibility premium; concrete nevers | B3 matrix ("Act" column = reversibility) · B5 envelope anatomy | 2, 3 |
| F4 ledgers rate highest | B2 principles (P-Records) · C1 records & ledgers | 1, 4 |
| F5 the number (7.5/wk; 30% zero) | B5 volume budget worked example · D1 open questions (OQ2) | 3 |
| F6 weights, defaults, signed objective | B2 principles (P-Values) · B3 planning rows · C1 deferral ledger | 1, 2, 4 |
| F7 is/ought gap | C5 getting started → maturing · A2 why this framework | 5 |
| F8 stable human core | B1 human roles · B2 all principles · B4 boundary test | 1, 2 |
| Zero abstract bans (F1/F3) | B3 matrix "never" rows written as concrete work classes | 2 |
| Trust moves in inches (F7) | A1 living framework + telemetry rationale; W5 launch narrative | 5 |
11. Caveats, exclusions and data hygiene (bind on every downstream use)
- Instruments evolved. W1 ran on Mentimeter with different option labels ("Make corrections with approval" ≈ act-with-approval) and a 7-option retained-rights list without stop authority or own-the-value-judgement. Cross-webinar tables in §1 and §9 are alignment reads, not identical instruments. The boundary vote is verbatim-stable from W2 onward; the decision-rights 0–5 poll is verbatim-stable across W2→W4.
- Small/late blocks. W3's boundary vote ran at C22 with n=18 after attrition (39 answered the opening question, 26 the closing one) — its percentages are directional. W4 fixed this by design (vote twice, C04+C21), which is itself a method finding for the paper.
- Unusable blocks — do not quote: W3
w3-s1-r4-untouchableword cloud (prompt misfired; 6 of 17 answers say they didn't understand); W3w3-monday-action(n=4). Both are flagged in [W3] and withheld from the public page. - Withheld data: W2's 11-section volunteer questionnaire drew 4 responses — below the MIN_N=10 privacy gate; its contents are deliberately not in the results file nor here.
- Paired beats unpaired. Where this document gives a paired figure, comms must use it (W4 ladder +0.44 n=25; W4 stop 48→64 n=25; W4 comfort +0.18 n=17; W3 trust +0.43 n=14). Unpaired shifts mix a change of mind with a change of room.
- Cross-room reads (W2 71% → W4 33/24% on the trigger instrument) compare different audiences four weeks apart; say "the community", not "participants changed their minds".
- Percentages reproduce each file's rounding; multi-selects sum >100%; ratios in §3 are computed from raw counts.
- Consent & anonymity: figures are consented responses; W3/W4 tokens are HMAC-anonymised and irreversible. No individual is identifiable; keep it that way in derived work.
- Sector mix differs per room (W3 transport/consulting-heavy; W1 government-heavy; W4 unmeasured). No claim in this document is sector-specific.
- Claim-language discipline for anything derived from this document: "the world's first Human-AI Boundary Framework for asset management"; cite PMI's June 2026 standard (world's first global standard for AI in project work) rather than competing with it; the Asset Zero claim keeps its "to our knowledge" qualifier. See
Phase 2 - AI Native Project/research/01-ai-native-pm-landscape.md§(d).
12. One paragraph for the paper (the spine in prose)
Across four webinars, June–July 2026, an Australian and New Zealand asset-management community (2,796 recorded responses) drew a consistent boundary: AI may analyse and recommend almost anywhere, may act only inside engineered, reversible, rate-limited envelopes behind human approval, and may never own purpose, values, risk acceptance or the stop. The ceiling held at analyse-or-recommend (67–72%) across data, predictions and money; it moved — once, by +0.44 rungs paired — when a room watched bounded, ledgered, stoppable action for forty minutes, and it moved to approval-gated action, not autonomy. The community priced approval above interruption for three sessions (1.1×→2.7×), then re-priced the stop the moment it saw a storm (48%→64% paired, with zero drops). It valued reversibility at 38 percentage points on the same work order. It rated ledgers — named, dated, signed records of deferrals and actions — above every other mechanism offered (4.07, 4.27 of 5). It drafted its first quantitative boundary parameter: a median of 7.5 machine-created work orders per planner per week, with a third saying never. And it reported an is/ought gap the framework exists to close: 36% could pause the machine within the hour today, 22% have drilled it, and the median practitioner guesses that 0% of AI-generated plans are audited. The boundary, on this evidence, is a governance choice — and this community has now specified it, decision by decision.
Provenance: AI-drafted (Stream 1, 2 Aug 2026) from the four results files listed in §0; verified figure-by-figure against those files; fact-check pass run before repo commit. Interim boundary registration: Recommend/draft-only — a human (Kai) reviews and commits. Status: PROVISIONAL until WG cycle 1 opens and a named human signs off.