Try Northstar · Episode C of three

What did the change accomplish?

Review the observed results and decide what should happen next. Look beyond the headline performance measure to customer outcomes, variation, cost, and operating consequences.

Use your AI to produce a concise recommendation. Review its reasoning, then compare it with the reference answer.

Northstar v1.2 is a fictional teaching case. All organizations, source records, figures, and events are authored for the exercise.

  1. What to fund
  2. Ready to begin?
  3. What the results justify

Start with two files and one prompt.

  1. Download these two files.

    The packet includes earlier evidence and teaching decisions, the post-release results, and the data tables. It withholds the reference value judgment. Prefer one ZIP? Download the learner packet.

  2. Attach both files to a new AI conversation.

    Use an AI that accepts Markdown attachments. Ask it to confirm that it can read both files. Keep the reference answers and later episodes out of this conversation.

    If you completed an earlier episode, paste your continuation note too. Reconcile your recommendation with the supplied teaching decision before beginning this stage.

  3. Copy this prompt and send it to your AI.

    Episode C exercise prompt

    Use the attached SCALE-AGENT-INSTRUCTIONS.md as the working method. EVIDENCE-PACKET.md contains fictional Northstar evidence, supplied decisions, and calculation data available by May 6, 2026. Use earlier material as background and answer the current results question below.
    
    The earlier decisions in the packet are scenario facts, separate from my recommendations. Preserve any earlier conclusion I supply as my previous work; explain changes without rewriting it. If I supply none, identify a conclusion from the packet's earlier evidence that merits reconsideration.
    
    Assess the day-90 results and recommend the next action. Produce a concise decision brief, about one to two pages with a small calculation table if useful, that:
    - Gives a direct recommendation and identifies the authority required.
    - Reconciles reporting periods, denominators, changed work scope, and branch variation.
    - Shows the arithmetic behind at least one material claim, citing the source, period, numerator, denominator, and scope.
    - Examines customer outcomes, cost, and operational conditions alongside the headline performance measure.
    - Distinguishes what the evidence establishes from uncertainty and competing explanations.
    - Compares a credible alternative, states conditions and stop or review triggers, and explains what would change the recommendation.
    - Identifies one earlier conclusion to narrow, reverse, or retain with a reason.
    
    Use the attached definitions and tables. Keep missing data visible; do not invent transaction records, causal attribution, approvals, or results.
    
    Finish by identifying the conclusion most vulnerable to challenge and the evidence that could change it. I will review the draft. Keep my review separate from a fictional business decision. A reviewed brief and a short correction note are sufficient; no JSON export or complete architecture model is required.
    Download prompt

Read the evidence on this page ↓

After your AI responds

Review the recommendation before comparing answers.

  • Open the sources behind the important claims. Does the evidence support them?
  • Challenge the strongest alternative, missing information, and anything that would change the recommendation.
  • Correct the draft and keep your brief with its uncertainty, conditions, and named decision maker.

If you want to continue later, ask the AI for a short continuation note preserving your corrections, rejected claims, unresolved questions, and next step. Save it with your brief and source files.

Have a recommendation you can explain?

Compare the Episode C answer →

The reference is one reasoned response, not the only acceptable conclusion. A different recommendation should explain its evidence and tradeoffs.

Evidence available at this decision

Read Episode C online.

The two-file setup’s evidence download combines everything needed for this episode. The reader below separates the current packet from any earlier context.

Earlier context · Episode A evidence

Northstar: evidence available before the investment decision

Learning edition: 1.2 · Decision date: 2026-01-09 · Classification: public, entirely fictional. Boundary: E-001–E-007 only. This packet contains no subsequent decision or result.

These are newly authored fictional source extracts for a teaching exercise. They extend Northstar v1.1 with additional uncertainty, alternatives, and financial assumptions; they are not recovered underlying records from v1.1 or observations from a real company. Quoted speakers and reviewers are fictional. Source authority establishes who supplied a claim, not that the claim is correct.

The request

Northstar Equipment Services has $780 million in annual revenue, 48 branches in twelve states, and 3,200 employees. It maintains and rents commercial equipment. Customers need equipment available when their own work requires it.

Priya Shah, EVP Service Operations (ORG-09), asks:

Build us a capability map and tell the Council whether the proposed scheduling platform fixes our service-promise problem.

Jordan Reed chairs the Investment Council (ORG-13), which controls funding. Avery Chen is the business architect (ORG-14); Avery can frame, analyze, recommend, and record, but cannot approve investment. Morgan Ellis, CFO (ORG-12), governs financial definitions.

Strategy records identify O-01: improve reliable service promises with appropriate branch discretion; O-02: protect margin; O-03: reduce customer disruption and repeat work. The executive target is 95% promise adherence by 2027. A possible first release would test feasibility, not establish the strategic target or enterprise repeatability.

Existing architecture fragments

The following identifiers are established; their relevance and relationships still need review. Do not invent a complete enterprise map.

ID Accepted name and boundary
VS-01 Restore Equipment Availability; value belongs to the customer whose equipment becomes usable
ST-02 Define service promise; move from accepted need to a understood, feasible commitment
C-110 Service Promise Management; govern the commitment, using resource signals without reserving parts or scheduling technicians itself
C-130 Resource Scheduling; match authorized work to qualified people, time, geography, and capacity
C-140 Parts Availability Management; determine, reserve, position, consume, release, and expire part availability
C-170 Service Performance Management; define and explain measures and guardrails, without acquiring Finance's authority
C-180 Business Information Governance; govern business meaning, ownership, quality, and semantic change
BO-02 Customer Promise; a commitment, not every ETA estimate; preserve original and revised versions
BO-05 Part Reservation; proposed owner ORG-08 VP Supply Operations; inventory record is a candidate source, not confirmed authority
BO-06 Technician Assignment; skill qualification and assignment acceptance are separate checks
I-01 Connected Service Promise; proposed change boundary, not an approved solution

ORG-02 branch managers control local capacity and authorized exceptions; ORG-05 Regional Operations resolves cross-branch constraints; ORG-07 Field Service owns field performance; ORG-08 Supply Operations owns parts performance; ORG-10 Information Governance confirms information authority. Architecture stewardship does not transfer operating accountability.

E-001

Source: Operations dashboard extract and definition notes, 2026-01-05; owner ORG-09; twelve-month reporting across all 48 branches. Directionally useful, not a reconciled pilot baseline.

Source fragment Recorded observation
E-001.1 enterprise dashboard Reported M-01 promise adherence: approximately 82%. Local denominators were combined without a common cancellation rule.
E-001.2 metropolitan definition Count completed work against the latest accepted appointment; remove customer cancellations.
E-001.3 rural definition Count work against the original promised window unless an operational manager approves a revision; cancellations remain until reviewed.
E-001.4 customer desk note Missed or changed promises generate complaints, but service-recovery cases are not consistently linked to the original commitment.

Analyst warning: these definitions may produce different rates for identical service. Reconcile original versus latest promise, eligible population, revisions, cancellations, and reporting period before treating 82% as comparable to a pilot result. No customer-restoration-time baseline has been approved.

E-002

Source: Regional quality sample, 2026-01-06; owner ORG-07; repeat-visit review.

E-002.1: The sample reports approximately 14% avoidable repeat visits (M-02), against an 8% strategic aspiration. E-002.2: Two regions dispute whether customer-requested additional work, diagnostic return visits, and unavailable parts count as preventable. E-002.3: The sampling frame and event-level coding require validation. The apparent gap cannot yet support an attributable savings calculation.

E-003

Source: Five branch workshops and dispatcher observation notes, 2026-01-07; collected by Avery Chen; convergent observations, not a controlled causal study.

Locator Fictional source extract
E-003.1 dispatcher, metro “I promise a date while the customer is on the line. The part screen says allocated; I call later to find out whether it is physically held.”
E-003.2 supply supervisor “An allocation is a planning quantity. A reservation can expire. A screen refresh is not a guarantee that the part is still there.”
E-003.3 rural manager “The dates are difficult because the qualified technician may be three hours away. Approval for an exception sometimes takes longer than making another plan.”
E-003.4 commercial manager “If we wait for every confirmation before quoting a date, some customers will leave. Later promises may improve the dashboard while making our service worse.”
E-003.5 technology lead “The existing system exposes a timestamp and reservation identifier. I have not established whether all branches update them consistently.”

Competing explanations include premature commitment, unreliable reservation meaning, constrained travel/skills, unusable exception authority, and measurement artifacts. Neither these accounts nor a capability diagram establishes their relative causal contribution.

E-004

Source: Vendor demonstration and internal integration note, 2026-01-08; technology owner ORG-11.

E-004.1: In a controlled demonstration the proposed scheduling product blocked commitments when required part and skill fields were absent. E-004.2: No live Northstar reservation feed, representative rural work, adoption test, or service-value experiment was used. E-004.3: The vendor proposes $4–7 million over 18–24 months, subject to scope and commercial agreement. This is not an accepted quote. E-004.4: An internal team proposes a narrower reservation/skill-status feed and dispatcher exception queue using existing systems. Its $60–120k and one-to-two temporary FTE planning range assumes usable reservation identifiers and source-state semantics. It would not solve technician scarcity, travel, or disputed promise policy.

The narrow option is credible enough to assess. It is not proven feasible by the availability of an API.

E-005

Source: Finance planning memorandum, 2026-01-08; Morgan Ellis, ORG-12.

E-005.1: Governed enterprise gross service margin (M-03) is 31%; branch and work-mix variation are material. A proposed guardrail is at least 31% cumulative at investment reviews. E-005.2: The case has no approved causal estimate of avoidable loss. For planning only, Finance brackets recoverable service credits/overtime associated with the candidate problem at $9–18k per month across the intended test boundary. The assumption is unvalidated, gross, and excludes lost demand; do not count it as an achieved benefit or an ROI. E-005.3: A 90-day delay therefore exposes a scenario of $27–54k in those costs if nothing else changes. Some delay may be necessary to prevent a larger wrong investment; the range does not settle the decision. E-005.4: The $150–300k broad operating-test range must include internal labor, external work, training, branch participation, reconciliation, and measurement. Report recurring manual burden separately from one-time change cost. Peak FTE is a capacity constraint, not total labor effort. E-005.5: Before material commitment, identify which uncertain fact would change the next decision, what cheaper work can resolve first, and the maximum amount worth spending to learn it. No $300k authorization should be treated as an instruction to spend $300k.

E-006

Source: Two-week commitment review, 2026-01-08; provisional M-04 sample.

E-006.1: 68 of 180 reviewed commitments had recorded part and skill confirmation before commitment: 37.8%, conventionally reported as 38%. E-006.2: The sample spans six candidate branches and broad work types. It is not the later three-branch eligible cohort and cannot be annualized into pilot volume. E-006.3: This checks part and skill signals, not every policy obligation. A commitment can pass those checks and still lack required customer communication, geography validation, or a proper approval. E-006.4: The review does not establish whether missing evidence means a missing action, a recording failure, or a genuinely infeasible commitment.

E-007

Source: Branch capacity and options worksheet, 2026-01-09; Regional Operations with Finance and Delivery; planning evidence.

Candidate test conditions: one metropolitan branch, one mixed-fleet branch with part constraints, and one rural branch with long travel and scarce skill substitution. Select deliberately for variation, not statistical representation. Verify leadership stability, safe participation capacity, auditable records, and competing change before activation.

Option Scope Planning cost / capacity What it might establish Material weakness
A0 Reconcile measures and improve reporting only $25–75k; 0.5–1 temporary FTE Baseline comparability and problem extent Does not directly change promise practices
A1 Three-branch operating experiment using existing systems $150–300k; 4–6 temporary FTE including measured branch participation Policy/exception feasibility, manual burden, differentiated branch constraints Requires scarce capacity; may not distinguish each cause
A2 Immediate twelve-branch operating expansion $800k–1.5m; 10–15 temporary FTE Broader variation sooner Commits before the core practice and measures are proven
A3 Procure and deploy enterprise platform $4–7m; 20–30 temporary FTE over 18–24 months Technology fit and operating results after major commitment High path dependence; present evidence proves functionality only
A4 Targeted reservation/skill feed plus exception queue $60–120k; 1–2 temporary FTE Whether a narrow automation removes a material delay Depends on unresolved source meaning; limited policy/capacity coverage

A1 and A4 may be complements or alternatives. There is no requirement to recommend A1.

E-007.1: A five-business-day diagnostic is possible within $15k, charged within any eventual A1 cap: reconcile a sample of reservation events; time actual checking; inspect original/revised/cancelled commitments; test exception authority; verify available branch capacity. E-007.2: A prospective staged A1 envelope would release up to $15k for that diagnostic, up to $75k cumulative for policy/source/readiness work, and at most $300k cumulative only after an operational-entry decision. These are proposed controls, not recorded approval. E-007.3: Potential guards include M-06 no more than 3% unexplained productivity deterioration; M-08 at most $300k and six peak temporary FTE; customer-escalation tolerance agreed before release; original-promise fulfillment and time-to-restored-availability monitored alongside M-01. E-007.4: Operational readiness, customer harm, and unclear source authority can stop a test even after analysis is sufficient for an investment decision.

Your task

Recommend the smallest defensible next commitment, including a competing interpretation you take seriously. Identify what would make A0, A4, a different experiment, or stopping preferable. State whether your architecture relationships are observed, inferred, proposed, or accepted. Give the Council a usable decision brief and preserve unresolved definitions.

Do not manufacture a probability of success, monetary value of information, or causation. If the proposed learning cannot change a decision, redesign it before requesting the full experiment.

Earlier context · Episode B evidence

Northstar: evidence available before readiness

Learning edition: 1.2 · Decision date: 2026-01-30 · Classification: public, entirely fictional. Boundary: previous Episode A decisions and E-008. No release-performance evidence is available.

These newly authored extracts extend the fictional v1.1 case. They are not actual meeting minutes or recovered enterprise records.

Prior decisions now available

Jordan Reed (ORG-13), after review with Morgan Ellis and Priya Shah, recorded DEC-001 on January 9: conditionally authorize A1/R1 within a $150–300k planning range, with a $15k diagnostic stage, $75k cumulative preparation limit, and $300k total cap. Activation requires accepted policy, definitions, owners/source statuses, branch selection, and readiness. The cap includes branch participation. No authority to expand to twelve branches follows.

DEC-002 defers enterprise-platform procurement. A4 targeted integration remains a comparator if residual gaps justify it. Procurement needs separate evidence and business-case authorization; neither a successful prototype nor adoption alone selects a vendor.

E-008

Source: signed preparation records and readiness review packet, 2026-01-30. Owners: Priya Shah ORG-09; Morgan Ellis ORG-12; Casey Ortiz ORG-15 readiness chair; Supply and Information Governance for source conditions. Authority: approved only for the stated R1 scope. Measures and source acceptances are versioned; operating performance remains unobserved.

E-008.1: diagnostic report

Avery Chen and Supply staff inspected 24 reservation-event sequences. Seven had delayed status updates; four uses of “allocated” did not denote a confirmed reservation. This purposive diagnostic establishes a semantic/control defect, not an enterprise failure rate. Timing 18 dispatcher checks found median checking effort of six minutes and a range of two to seventeen minutes; sample selection and observer effects limit extrapolation.

A4's feed is technically plausible, but directly publishing the present status would carry the ambiguity into a faster interface. One branch already uses a reliable confirmation call; rural technician coverage and exception latency remain unresolved. These observations justify investigating an operating practice while preserving targeted integration as a rival. They do not establish that software is unnecessary.

The Council's January 19 stage-release addendum to DEC-001, following review of the diagnostic by Priya Shah and Morgan Ellis, authorized preparation spending up to $75k cumulatively. Neither the initial envelope nor completion of a checklist released those funds automatically.

Stage-0 diagnostic cost: $12k. Preparation expenditure after diagnostic: $58k. Cumulative $70k remains below the $75k preparation limit. The updated forecast is $286k total and no more than six peak temporary FTE. Finance records these as forecasts, not results.

E-008.2: policy and rights

POL-01 v1.0 requires valid BO-02 promise data and part, skill, geography, and capacity confirmation for the normal path. Where confirmation is unavailable, a complete approved exception identifies reason, approver, expiry, and customer communication.

ORG-02 branch managers approve delegated local exceptions; ORG-05 Regional Operations handles cross-branch cases. Dispatchers apply rules but cannot waive guardrails. A changed promise retains the original and its version links; a broken promise opens customer recovery before closure.

ORG-15 may activate, stagger, adapt, pause, or stop R1 within its delegated scope. For the M-03 margin guardrail only, after CFO validation it may authorize one contained remediation interval to the next formal review. It cannot renew that exception, exceed the investment/capacity cap, activate R2, or reopen procurement.

E-008.3: object-source decision

ORG-08 is confirmed as the BO-05 business-object owner. The inventory reservation record is accepted as a provisional R1 status source, subject to twice-daily reconciliation. It is not authoritative enterprise-wide.

BO-02's existing record is accepted for R1 with versioning, original/revised linkage, and audit rules. Acceptance does not declare the enterprise future-state technology.

OI-02 is closed for its R1 question. OI-06 opens under ORG-08 and ORG-10: determine enterprise/future-state BO-05 authority before R2 activation. The boundary is material; replacing “provisional R1” with “authoritative” changes the decision.

E-008.4: branch and evidence design

Selected branches: metropolitan, mixed-fleet, and rural. Minimum 90-day volume: 100 eligible promises per branch and 450 overall. This is an operational sufficiency threshold, not a statistical power calculation or guarantee of representativeness.

The January 3–February 1 baseline window is prospective and not yet complete on January 30. Eligibility is fixed to specified test work types; retain cancelled, revised, declined, and removed-scope counts separately. A source may not delete unfavorable observations to meet a gate. Record both original and latest accepted promise performance and hours until equipment is restored. Before activation, Finance and Operations must confirm the finished baseline reconciliation.

M-04 checks recorded part/skill signals. M-07 checks complete compliant workflow, including approved exceptions. Neither is universally a subset of the other. The approved reporting plan uses day-30 and later-period operational snapshots for these measures; labels must identify their periods. Do not compare a current-practice percentage with a cumulative result as if denominators were identical.

E-008.5: guards and measures

Measure R1 interpretation / action
M-01 92% cumulative latest-promise adherence is an interim feasibility gate; 95% remains the strategic target
M-02 Repeat visits remain provisional; rebaseline cause coding before R2
M-03 At least 31% cumulative margin at investment review; CFO validates; no expansion on breach
M-04 90% required part/skill confirmation in the later operational review period
M-05 Day-30 escalation rate at most 0.5 per 100 above the approved baseline; severe harm pauses affected work
M-06 No more than 3% unexplained technician-productivity deterioration; Field Service and Finance validate; no scale while unresolved
M-07 At least 85% compliant use by day 30 and 95% in the later operational period
M-08 At most $300k total and six peak temporary FTE; forecast breach requires scope/funding action before excess occurs
POL-01 At least 90% of exception requests resolved within one operating shift; repair authority/capacity before enforcing or scaling
Learning-edition balancing view Report original-promise fulfillment, revisions, cancellations, declined work, and restoration hours; adverse movement triggers customer-value investigation even if M-01 passes

E-008.6: readiness observations

Area Evidence currently available Unresolved condition
Ownership/policy Rights signed; tabletop completed None within R1 scope
Information BO-02 linkage passed; one branch has delayed BO-05 updates Twice-daily reconciliation with named Supply duty owner
Exception capacity Peak-volume simulation exposes rural approval backlog Assign and test regional backup approver
Training 91% complete; one shift has not passed scenarios No activation for an untrained shift
Measures Definitions agreed; parallel-run tests reconcile Finish the January baseline before activation; preserve unresolved M-02 coding
Recovery Rollback and record-retention procedure tested Maintain original promises and recovery ownership on rollback

The forum has not yet made DEC-003. Write a decision with explicit conditions and expiry; do not infer a waiver from the investment authorization.

Northstar: evidence available for the day-90 decision

Learning edition: 1.2 · Decision date: 2026-05-06 · Classification: public, entirely fictional. Boundary: E-009/E-010 and the earlier dispositions below. DEC-005 has not yet been made.

All source records, counts, names, and decisions in this exercise are fictional. The exact-count datasets and original-promise/restoration outcomes are newly authored v1.2 extensions; they were not recovered from v1.1. The additions intentionally make the next decision less comfortable than reading the aggregate dashboard alone.

Decisions already made

DEC-001 conditionally authorized the three-branch A1/R1 test within $150–300k, six peak temporary FTE, and staged commitment. DEC-002 deferred enterprise procurement.

DEC-003, recorded by Casey Ortiz (ORG-15) on January 30, authorized a staggered February 2 start only after the remaining shift passed training and the January baseline was reconciled. It accepted provisional BO-05 reconciliation and a regional backup approver until the day-30 review. These conditions conveyed no R2 or procurement authority.

The readiness addendum dated February 1 records training completion and baseline reconciliation. ORG-08 owns BO-05; its R1 source remains provisional. OI-06 for enterprise source authority remains open.

E-009

Source: day-30 operating/Finance extract and review, 2026-03-03. Owners: Priya Shah ORG-09; Morgan Ellis ORG-12; ORG-07 Field Service; Delivery Management. Period: February 2–March 3 inclusive, 30 days.

Locator Observation
E-009.1 128 of 144 eligible promises met the latest accepted promise: 88.9%, reported as 89%
E-009.2 121/144 had part/skill confirmation (M-04 84%); 112/144 followed complete policy or an approved exception (M-07 78%)
E-009.3 M-03 margin 30.8%; M-06 technician productivity 4.1% below comparable baseline
E-009.4 M-05 escalations 6.5 per 100, compared with approved pre-release 6.2: within the 0.5 tolerance
E-009.5 Repeat visits 12% under still-disputed coding; one-shift exceptions 76%; rural approval backlog persists
E-009.6 $118k spent; forecast $340k total; peak capacity 5.8 FTE; forecast exceeds the authorized cap
E-009.7 Source lag creates reconciliation work. No comparison yet establishes that the proposed enterprise platform is the cheapest effective remedy.

Known day-30 decision: DEC-004

After CFO validation of margin/cost, Field Service validation of productivity, and Delivery validation of the forecast, Casey Ortiz used the one permitted remediation interval to continue R1 until the day-90 formal review.

The decision prohibited scope expansion, imposed the M-06 no-scale response, and closed a low-volume work type to new pilot admissions from March 4 to avoid forecast overspend. It required weekly margin/workload review, correction of BO-05 reconciliation, and a tested regional backup approver. Severe customer harm could pause a branch; extension of the margin exception or any cap increase required the Investment Council.

Completed observations from the removed work type were retained. Removing future scope was not permission to erase past failures. The scope change affects comparison and must be shown to the next forum.

E-010

Source: day-90 extract, branch segmentation, audit, and Finance review, 2026-05-02. Review meeting: May 6. Owners as above, with customer-outcome review by ORG-06. Detailed records: see datasets/README.md and the CSV files in this folder.

E-010.1: operational headline and reporting periods

There are 486 cumulative eligible promises: 176 metropolitan, 166 mixed-fleet, 144 rural. Of these, 167, 156, and 128 respectively meet the latest accepted promise. Total 451/486 = 92.8%, reported as 93%; the branch rates round to 95%, 94%, and 89%. Each branch passes 100 observations and the total passes 450. These counts meet a volume rule, not a representativeness test.

Later-period M-04 is 315/342 = 92.1% and M-07 is 328/342 = 95.9%, reported as 92% and 96%. Their period is days 31–90, not the full cumulative 90 days. Early M-07 noncompliance cannot disappear from a cumulative denominator.

M-04 records part/skill confirmation only. M-07 records all required policy steps or a complete approved exception. A promise can have confirmation but incomplete documentation; a fully approved exception can be compliant without normal confirmation. These measures are not nested sets.

E-010.2: cost and guardrails

Finance reports cumulative gross service margin of 31.2%. Technician productivity is 2.0% below the comparable baseline, within its 3% guardrail. Total change cost is $287k and peak capacity is 5.8 FTE, within the authorized band. Cost includes $12k diagnostic, $58k preparation, and $217k release/measurement/change activity; ongoing manual-control cost requires a separate operating estimate.

Escalations are 5.6 per 100. Repeat visits are 11%, still above the 8% strategic aspiration and still subject to coding reconciliation. One-shift exception completion is 88%, below the 90% criterion. Rural adherence remains 89%. These are unresolved limitations, not noise to average away.

E-010.3: scope and stable-cohort comparison

The first 30 days contain 114 retained-scope promises with 106 latest-promise successes and 30 subsequently closed-scope promises with 22 successes. Days 31–90 contain 342 retained-scope promises with 323 successes.

Thus the retained-scope latest-promise rate moves from 93.0% to 94.4%, a 1.5 percentage-point change. The broader aggregate moves from 88.9% early to 94.4% later, partly because the poorer-performing work type no longer enters. All 30 earlier closed-scope records remain in the 486 cumulative total.

Do not attribute the whole aggregate movement to policy adoption. Time, case mix, selection, changes in demand, and the intervention may all contribute.

E-010.4: added customer-value evidence

The v1.2 customer-outcome extract reports original-promise fulfillment of 120/144 early and 268/342 later. Within retained scope it falls from 100/114 (87.7%) to 268/342 (78.4%). Mean restoration time in that retained cohort rises from 40 to 48 hours.

These are descriptive comparisons, without case-mix adjustment or a control group. They do not prove that the new policy caused harm. They do establish that latest-promise adherence alone is insufficient evidence of improved customer value. Examine promise revisions, demand refused or diverted, operating mix, and capacity before declaring success.

The separate cancellation/decline denominators in the dataset keep changes in the admitted population visible; they must not be folded into fulfilled-promise percentages without a defined question.

E-010.5: request to the Council

The sponsor proposes designing R2 because the operational feasibility gate is met. Customer Operations asks whether later accepted dates conceal worse service to the original need. Finance asks for the recurring manual-control cost and a comparison of A4 targeted integration, further operating correction, and platform alternatives. No one has a controlled attribution estimate.

Avery must recommend the next decision. The Council may authorize design/readiness, restrict or stop portions of the practice, require investigation, or decline expansion. No existing decision authorizes twelve-branch activation or procurement.

Episode C · Inspect the numbers

Download the underlying case tables.

All six tables and their definitions are already embedded in your EVIDENCE-PACKET.md. No additional uploads are needed. These separate CSV files are optional if you want to check the calculations in a spreadsheet.

Download optional spreadsheet calculation data

DatasetUse
Baseline scopeInspect the comparison boundary and baseline context.
Branch resultsReconcile branch counts and cumulative performance.
Control measuresDistinguish the confirmation and policy-compliance measures.
Customer outcomesCompare original commitments and restoration time.
Demand decisionsInspect admissions, declines, and missing cancellation information.
Scope bridgeKeep removed work in the historical denominator and compare retained scope.
Read dataset definitions and calculation guidance

Northstar learning edition v1.2: Episode C datasets

Authored fictional teaching data. These are deliberately constructed aggregate counts for the Northstar learning edition. They extend the v1.1 narrative; they are not recovered original observations, real company records, or empirical evidence that SCALE or AI improves outcomes. No invented row is presented as a real person's service history.

The v1.1 day-90 branch volumes and rounded adherence percentages are preserved. Exact adherence counts are newly authored here to make the arithmetic inspectable: 167 of 176 metropolitan, 156 of 166 mixed, and 128 of 144 rural. These sum to 451 of 486 (92.7984%, displayed as 93%). The exact counts do not purport to reconstruct missing v1.1 data. Customer restoration, original-promise, admission-decision, and denominator-bridge observations are newly authored additions, with deliberately mixed implications.

Files and boundaries

File Unit represented by one row Use
branch-results.csv One branch across all 90 days Reproduce the headline and branch rates
scope-bridge.csv Branch x completion window x work-type scope Reconcile early, later, retained, and removed scope
customer-outcomes.csv The same completed restoration cohort as the scope bridge Calculate original-commitment fulfillment and mean elapsed restoration time
control-measures.csv All completed promises within one window Compare separate confirmation and workflow-policy checks
demand-decisions.csv Admission decisions issued within one window Inspect declined requests outside the completed-promise denominator
baseline-scope.csv One distinctly bounded pre-pilot evidence set Prevent inappropriate baseline comparisons

A shared source ID identifies an authored source extract, not a citation to an external dataset. E-009 is the day-30 extract dated March 3, 2026; E-010 is the day-90 extract dated May 2, 2026. E-006 identifies the provisional diagnostic-sample counts and scope. E-008 includes the pre-release baseline finalized in its February 1 addendum; that completed baseline is not available to the January 30 readiness recommendation. E-001's legacy enterprise percentage is retained only as narrative context: missing numerator, denominator, and dates are blank, not zero.

The pilot runs February 2 through May 2, 2026, inclusive (90 calendar days). Days 1–30 end March 3; days 31–90 run March 4 through May 2. The review is May 6. Dates describe this fictional case, not publication or retrieval dates.

Counting rules

  1. An eligible completed promise is one completed equipment-restoration episode within the permitted work-type scope of one of the three pilot branches. There is one promise record per completed restoration episode in this authored dataset. Completion date determines the window.
  2. Latest-promise adherence counts completion by the most recent approved commitment recorded and communicated before completion. It can improve when promised dates are revised; the separate original-promise measure deliberately retains that exposure. Aggregate counts cannot establish whether every change was timely or appropriate.
  3. Original-promise fulfillment counts completion by the first communicated commitment, without replacing it with a revised date. The same completed-episode denominator is used for both promise measures.
  4. Elapsed restoration hours run from recorded equipment-unavailable time to recorded restoration of usable availability. Calendar hours include nights and weekends; this is neither labor time nor billable time. Mean hours = sum of elapsed restoration hours / completed restoration episodes.
  5. Confirmation passes only when both a part reservation and a qualified-skill reservation are corroborated at the recorded commitment check. Estimated availability is insufficient.
  6. Whole-policy compliance independently checks required workflow, communication, documentation, exception routing, and approvals. The fictional policy permits an explicitly communicated conditional commitment through its approved exception route. Such a record can comply with policy while failing the part-and-skill confirmation test. Conversely, a confirmed part and skill do not guarantee correct communication or exception handling. The two pass counts are not nested sets and cannot be subtracted to infer an overlap.
  7. Admission decisions are counted by decision date for all requested work at the selected three branches, including the work type removed from pilot scope. Declines never count as fulfilled promises. The cancellations_after_admission field is intentionally blank because no reconciled cancellation count was supplied; blank is not zero. These are different event cohorts from completions, so admission counts must not be joined to completion counts as if they were the same records.
  8. CSV percentage columns are calculated from their adjacent numerator and denominator and rounded to four decimals. For aggregate rates, sum counts first; do not average rounded branch rates. Blank fields mean unavailable.
  9. Summary rows are not duplicated in the CSVs. Summing disjoint scope-bridge rows produces the branch-results population. Do not sum branch-results and scope-bridge together.

Scope change and denominator bridge

At the day-30 decision, one low-volume work type leaves pilot scope. In this authored scenario all 30 admitted episodes of that type have completed by March 3. Subsequent requests remain visible in admission decisions but produce no further pilot completions of that type. The 30 completed episodes remain in every cumulative 90-day count: there is no retrospective deletion.

Completion window and scope Completed promises Latest commitment met Rate
Days 1–30, retained work types 114 106 92.98%
Days 1–30, subsequently removed type 30 22 73.33%
Days 1–30, all eligible scope 144 128 88.89%
Days 31–90, retained work types 342 323 94.44%
Days 1–90, all eligible completed scope 486 451 92.80%

The apparent 89% to 93% headline change mixes an early-window measure with a cumulative measure and a changed work-type composition. A retained-scope comparison is 106/114 to 323/342. That controls the named scope exclusion, but it does not hold individual jobs, severity, branch weights, seasonality, or other case mix constant. “Stable scope” is not a matched-person or matched-job cohort, nor a causal estimate.

Day-30 confirmation and policy figures are 121/144 = 84.03% and 112/144 = 77.78%. The day-90 reporting pack's 92% confirmation and 96% policy values use the later window, 315/342 = 92.11% and 328/342 = 95.91%. They are not cumulative day-90 percentages. For comparison, all-90-day counts would be 436/486 = 89.71% confirmed and 440/486 = 90.53% policy-compliant. A label such as “day-90 result” must therefore identify its actual measurement window.

Customer outcome counterchecks

Measure Days 1–30 Days 31–90
Original commitments fulfilled, all completed scope 120/144 = 83.33% 268/342 = 78.36%
Original commitments fulfilled, retained work types 100/114 = 87.72% 268/342 = 78.36%
Mean restoration time, all completed scope 5,760h/144 = 40h 16,416h/342 = 48h
Mean restoration time, retained work types 4,560h/114 = 40h 16,416h/342 = 48h
Declined requests, all admission decisions 12/180 = 6.67% 42/420 = 10.00%

Restoration time also increases within each retained branch slice: metropolitan 34h to 42h; mixed 40h to 46h; rural 46.67h to 57.8h. The deterioration cannot be explained solely by dropping the low-volume work type or changing branch weights. It still cannot be attributed causally to the pilot without investigating other changes.

The discrepancy between latest-commitment adherence and original commitments is a question for analysis. It is not proof of intentional manipulation. Do not infer customer value solely from the headline adherence improvement.

These are completed-case observations. Outstanding work, cancellations after admission, customer severity, promise revision timestamps, distributions, and reasons for declining requests are not supplied. There is no median, percentile, matched control, or statistical uncertainty estimate to calculate from the available totals. Ask for those records before claiming population-wide service improvement or causation. Do not invent individual event rows from these aggregates.

Baselines are separate evidence sets

  • The historical 82% enterprise report covers 48 branches over a 12-month reporting period. Its missing counts are not reverse-engineered.
  • The provisional two-week diagnostic examines 180 commitments from six candidate branches and broader work types; 68/180 = 37.78% meet the part-and-skill check. It is not a two-week rate for the three-branch pilot and is not comparable to adherence.
  • The approved 30-day pre-release baseline covers the selected three branches and retained work types, January 3 through February 1, 2026: 98/120 = 81.67% latest-promise adherence. It supports a scoped descriptive comparison, not causal attribution or volume extrapolation.

Reproducibility and limits

All sums and rates can be recalculated from these CSVs. These checks establish internal arithmetic consistency only. The package deliberately supplies aggregate evidence, not a transaction-level audit trail; it cannot independently verify timestamps, duplicated records, eligibility coding, or the stated counting rules. An analyst should distinguish “recalculated from authored counts” from “validated against source transactions.”

Keep provenance with any chart or brief: Northstar v1.2, authored fictional teaching data; source file; scope; numerator/denominator; measurement dates. No customer-value, return-on-investment, repeatability, procurement, or AI-productivity claim is established by arithmetic alone.