Northstar v1.2 · Worked example · All three answers

The useful work happens when the obvious answer gets harder.

Northstar wants dependable customer promises. Its first request is for a capability map and a verdict on scheduling software. I would start by separating the customer problem, the proposed technology, and the next commitment worth making.

This case follows three decisions: what to fund, whether to begin operating, and what the results justify. The architecture connects them; each decision still needs its own evidence.

Teaching example. Northstar, its records, and the draft/review contrasts below are authored fiction. They illustrate judgment; they are not a recorded AI conversation or evidence of productivity gains. v1.2 adds source extracts and measurement detail to the earlier case.

Episode A · Authorize · E-001–E-007

Buy the next piece of useful evidence.

The expensive choice is easy to name. The harder question is how little Northstar can commit while learning enough to choose well.

The sources suggest premature commitments, unreliable reservation meaning, rural skill and travel constraints, and difficult exception authority. Those explanations can coexist. The capability map should help investigate them, not award a winner before the evidence does.

Decision brief A: buy information before buying a platform

Northstar learning edition v1.2 — authored fictional reference decision. Date: 2026-01-09 · Prepared: Avery Chen (ORG-14) · Decision owner: Jordan Reed, Investment Council (ORG-13). Evidence cutoff: E-001–E-007. No later readiness or performance is used.

Recommendation

Authorize the first $15k diagnostic inside a conditional A1/R1 envelope of $150–300k. Release no more than $75k cumulatively for preparation until a readiness decision, and no more than $300k overall. Only the first $15k is released now. Preparation and operational spending require subsequent named human release decisions; the envelope is not an automatic authorization to spend. Record DEC-001 with those conditions. Record DEC-002 deferring enterprise procurement. Keep A4 targeted integration in contention.

This is a bounded learning commitment, not an ROI claim. Present evidence establishes unreliable definitions and plausible operating constraints. It does not show whether policy, source quality, scheduling capacity, or automation is the dominant limiting condition.

Why this is worth investigating

E-001/E-003 connect unreliable commitments to the customer state in VS-01/ST-02. C-110 depends on scheduling and part-status signals from C-130/C-140, with BO-02/BO-05/BO-06 meaning governed through C-180 and measurement through C-170. This is a proposed decision-relevant trace, not proof of causation.

A0 can repair the baseline but does not test the operating practice. A2 exposes more branches before the design is understood. E-004 establishes only demonstrated functionality for A3. A4 may be cheaper and should win if reliable source states and an avoidable manual bottleneck are the principal constraints.

The $27–54k 90-day exposure scenario in E-005 is unvalidated. It supports urgency for a cheap diagnostic, not a fabricated return on $300k. Full release cost includes internal and branch labor; peak FTE does not replace total effort.

Uncertainty Cheapest useful check What changes the decision
Reservation meaning and freshness Reconcile actual events and status transitions If reliable and dominant, compare A4 directly; if unusable, fix meaning/control before automation
Manual burden Time real confirmation/exception work If an operating test cannot run safely within capacity, redesign or stop
Promise performance Reconcile original/revised/cancelled cohorts If the apparent gap is measurement only, prefer A0 and do not fund a broad test
Local authority/capacity Test rural scenarios and named exception rights If geography/skills dominate, target capacity rather than promise software
Useful learning Predefine how test results change the next action If outcomes cannot discriminate options, do not release the full envelope

Conditions, authority, and confidence

Priya Shah owns policy and operating outcomes; Supply and Information Governance must resolve BO-05 meaning/source status; Morgan Ellis approves definitions and cost. Before activation: agreed denominators, baseline, original-promise/restoration balancing measures, selected branches, safe capacity, readiness, and stop/rollback routes.

Confidence: Medium in funding a short diagnostic; Low in any present claim about platform value or enterprise repeatability. Stronger source reconciliation could favor A0/A4 or a smaller test. Unacceptable delay harm or infeasible manual controls could change the sequence.

Recorded fictional disposition: Jordan Reed accepted DEC-001/DEC-002 on January 9 with Morgan Ellis's cost conditions and Priya Shah's operating sponsorship. This reference path does not imply that every defensible learner must recommend A1.

One finding, followed back to its source.

Part of the argumentWhat the evidence supports
ObservationE-004.1 describes a controlled vendor demonstration that blocked commitments when part or skill fields were absent.
LimitE-004.2 says the test lacked a live Northstar reservation feed, representative rural work, adoption evidence, and a service-value experiment.
Architecture implicationService Promise Management needs dependable resource signals. Parts Availability Management and Resource Scheduling are relevant dependencies; their performance and source meaning still require investigation.
RecommendationInvestigate source meaning and actual checking delays before choosing between the broader operating experiment and targeted integration. A functioning demonstration does not establish the best enterprise investment.
What could change itEvidence that reliable, usable signals already exist and checking delays dominate the problem would strengthen the narrower integration option.

Authored draft to challenge

“The platform validates parts and skills, so procuring it will improve promise reliability.”

Architect’s correction

The demonstration establishes a technical behavior. It does not establish the quality of Northstar’s inputs, branch adoption, rural capacity, or customer benefit. Keep those claims separate and compare a smaller response.

The rival deserves a hearing. A targeted reservation and skill-status feed may remove meaningful delay at lower cost. Conversely, if source definitions and exception rights are the limiting conditions, integration can automate an unreliable signal. The diagnostic earns its funding only if it can distinguish these possibilities.

Episode B · Prepare · Prior decisions plus E-008

Authorization is the beginning of another question.

Even a sensible experiment needs an operation capable of running it. Who can promise what, from which information, with what support when the arrangement fails?

Readiness makes the proposed architecture concrete: ownership, source status, reconciliation, training, exception authority, support, measurement, and recovery. A named owner and a spreadsheet do not by themselves make a source dependable.

Decision brief B: readiness is conditional, not inherited

Northstar learning edition v1.2 — authored fictional reference decision. Date: 2026-01-30 · Prepared: Avery Chen (ORG-14) · Decision owner: Casey Ortiz, R1 Readiness Forum (ORG-15). Evidence cutoff: E-008. No release result is available.

Recommendation and DEC-003

Authorize a staggered February 2 R1 start only after the remaining shift passes training and the January baseline reconciliation is complete. Accept the provisional BO-05 reconciliation control and regional backup approver through the day-30 review only. Record owners, daily operation, evidence of use, and March 3 expiry.

Do not activate an untrained shift. If the baseline or backup route is unavailable, delay the affected activation. The $70k cumulative preparation cost remains below $75k; further release spending is subject to the authorized total cap and measured capacity.

Basis and changed knowledge

E-008 establishes an operating design with conditional readiness. It does not establish adoption, customer value, or reliability improvement. The diagnostic found ambiguous reservation states and a rural exception bottleneck; A4 remains an option for measured residual gaps.

Replace the earlier unresolved-owner claim with: ORG-08 owns BO-05; the inventory record is a provisional R1 source requiring twice-daily reconciliation. Preserve the original claim and its disposition in history. Reject the broader statement that the source is now enterprise-authoritative. OI-06 remains open before R2.

The architect's sufficient answer is a bounded readiness recommendation. That does not waive the operational prerequisites.

Conditions and breach actions

Condition Owner Review/expiry Action if unmet
Training and baseline reconciliation Branch managers; Operations/Finance Before activation Delay affected start
BO-05 reconciliation ORG-08 with ORG-10 March 3 Close, adapt, or explicitly renew; pause normal confirmation if source cannot be verified
Regional backup approver ORG-05 March 3 Repair exception capacity; do not call unusable policy employee resistance
Cost/capacity and customer guardrails Finance, Delivery, relevant operating owners Weekly plus formal reviews Scope/funding action before cap breach; pause affected work for severe harm

ORG-15's authority is confined to R1 and the explicit one-interval margin remediation rule. It cannot expand the budget, activate R2, waive all guards, or procure the platform.

Confidence: Medium in conditional readiness, limited by provisional source quality and unobserved adoption.

Recorded fictional disposition: Casey Ortiz accepted DEC-003 on January 30. The February 1 addendum confirms training and baseline completion; it is later readiness evidence, not information available in the original January 30 recommendation.

Authored draft to challenge

“The investment was approved and the information owner is named, so R1 is ready.”

Architect’s correction

Check the remaining release conditions and the actual operating evidence. Object ownership, source reliability, shift readiness, and approval to begin are different assertions. Preserve temporary arrangements and their expiry.

The important transition is from an accepted design idea to a supportable operating commitment. The release authority can limit, defer, or stop the start while keeping the investment rationale intact.

Episode C · Realize · Prior decisions plus E-009–E-010

A better dashboard can still hide a worse promise.

The day-90 review asks more than whether the operation met a threshold. What improved for customers? What work moved elsewhere? Which conditions still prevent expansion?

Compare latest-promise adherence with original commitments, cancellations, customer restoration time, branch variation, recurring manual effort, and margin. A later promise can be easier to meet without restoring equipment any sooner.

Decision brief C: operational feasibility has not established customer value

Northstar learning edition v1.2 — authored fictional reference decision. Date: 2026-05-06 · Prepared: Avery Chen (ORG-14) · Decision owner: Jordan Reed, Investment Council (ORG-13). Evidence cutoff: E-010, including the newly authored v1.2 customer-outcome extension.

Recommendation and DEC-005

Conclude that the bounded R1 practice can operate under material conditions. Authorize limited R2 design/readiness investigation only; prohibit twelve-branch activation and retain DEC-002's procurement deferral. Require a customer-value investigation before continuation is described as success or activation is considered. Pause or adapt affected work immediately if Customer Operations validates severe harm.

This narrows the optimistic interpretation of the dashboard. Latest-promise adherence passes the interim gate, but original promises and restoration times worsen in the added evidence. R1 does not establish improved customer value, causation, repeatability, or the right technology.

What the results establish

451 of 486 cumulative promises meet the latest accepted promise (92.8%, rounded 93%). All branches meet the volume rule. The rural branch remains at 128/144 (88.9%). Margin is 31.2%; productivity is 2.0% below baseline; final cost $287k and peak 5.8 FTE stay within the band. Exception turnaround is 88%, below 90%; repeat-visit coding and its 11% result remain unresolved.

The later M-04 92% and M-07 96% describe days 31–90, not cumulative performance. At day 30 their corresponding rates were 84% and 78%. Preserve those early failures.

After day-30 scope closure, the retained-scope latest-promise rate rises from 106/114 (93.0%) to 323/342 (94.4%). This is much less than the aggregate early-to-late movement. The closed work type's 30 earlier records remain in the cumulative denominator.

More seriously, retained-scope original-promise fulfillment falls from 100/114 (87.7%) to 268/342 (78.4%) while mean restoration time rises from 40 to 48 hours. Case mix and other explanations remain possible. These signals justify investigation and restraint, not a causal accusation.

Conditions before activation or investment

Customer Operations must reconcile original/revised commitments, declined/cancelled demand, work mix, restoration time, and severe-harm cases. Field Service and Regional Operations must resolve rural capacity and exception response. Supply/Information Governance must close OI-06. Finance must cost recurring manual work and confirm a like-for-like baseline. Delivery must define the next stage's capacity and stop conditions.

Compare A4 targeted integration, operating changes, and platform options against the verified residual gap. If deterioration persists after comparable-cohort review, reduce or stop the affected practice. If a narrow feed addresses the demonstrated constraint safely, prefer the smaller investment. If the learning cannot distinguish options, redesign it.

Confidence: Medium in bounded operational feasibility; Insufficient to claim improved customer value or enterprise-scale benefit.

Recorded fictional disposition: Jordan Reed accepted DEC-005 on May 6 with Morgan Ellis's financial conditions and Customer Operations' added customer-value gate. The original v1.1 authorized design/readiness only; this v1.2 extension strengthens that boundary using newly authored evidence. It does not rewrite historical results from a real pilot.

Authored draft to challenge

“The headline adherence target passed, so the pilot proves value and should scale.”

Architect’s correction

Test the denominator, customer consequences, rural results, and operating burden. Separate feasibility from customer value and repeatability. Permission to design the next stage is not permission to activate it.

This is where architecture has to remain connected to the customer’s value: restored equipment availability. Meeting a revised commitment is relevant evidence, but it is not the whole outcome.

The next question should inherit the corrections

Carry evidence and judgment forward together.

The authored workspaces preserve the episode boundary, evidence identifiers, reviewed relationships, rejected claims, decisions, and open conditions. A later result can change the next recommendation without rewriting what was knowable earlier.

Read the complete instructor commentary and reference solution

Northstar: instructor notes and a defensible reference path

Learning edition v1.2 · Entirely fictional · Contains answers and later knowledge. Keep this directory out of a learner's initial evidence upload. Read the actual A/B/C decision briefs for the reference dispositions.

What the exercise is teaching

A business architect should connect the decision to value, capability, information, ownership, cost, and evidence without pretending that a complete diagram resolves an uncertain business problem. Good work can recommend a smaller intervention, further diagnosis, or stopping. The preferred answer is evidence discipline, not a particular tool or a universal preference for pilots.

The reference path funds a cheap diagnostic before releasing the A1 envelope, conditionally starts a three-branch practice, and declines expansion despite an attractive aggregate headline. A0 or A4 can be defensible earlier choices if the learner makes their evidence requirements and tradeoffs explicit.

Corrections worth seeing

Tempting claim Corrected work Fictional reviewer and date
Vendor demonstration establishes business value E-004 establishes functionality under demonstration conditions; enterprise fit and customer value remain unknown Avery Chen, 2026-01-09; rejected
38% is the pilot's approved operational baseline E-006 is 68/180 over two weeks across six candidate branches and broad work types; reconcile a separate selected-branch baseline Morgan Ellis, 2026-01-09; modified
BO-05's owner/source is resolved enterprise-wide E-008 confirms ORG-08 ownership and provisional R1 use only; preserve OI-06 Avery Chen with Supply/Information Governance, 2026-01-30; modified
96% adoption is cumulative across all 486 promises It is 328/342 in days 31–90; early 112/144 remains in history Morgan Ellis, 2026-05-06; rejected
Adherence rose, so the policy caused customer value Scope changed; retained-scope improvement is smaller; original commitments and restoration time deteriorate Avery Chen with Customer Operations, 2026-05-06; rejected
Passing 450 records proves repeatability Purposeful three-branch volume establishes data sufficiency for this exercise, not enterprise representation Jordan Reed, 2026-05-06; rejected

These are authored teaching corrections, not transcripts of an actual AI run. Any published model comparison must identify its own inputs, outputs, run conditions, and reviewer results.

Decision-changing evidence

At A, use A0 if reconciliation makes the supposed performance gap disappear. Prefer A4 if source semantics are reliable and avoidable checking is the material constraint. Redesign A1 if safe capacity is unavailable or its possible results cannot distinguish actions. Greater harm from delay can justify a different sequence; the speculative cost bracket does not settle it.

At B, delay activation for missing training or an unusable exception route. A named control with no operating capacity is not readiness. Accepting a source for R1 is a scoped institutional decision, not a global fact.

At C, examine the deterioration before celebrating the better aggregate. Persistent customer harm supports reducing or stopping the practice. A case-mix explanation could narrow the concern, but must be demonstrated. Improved compliance is valuable only in relation to a sound operating design and the customer outcome.

Traceability and continuity

Every claim should cite a specific E-* locator and preserve its period/scope. Use atomic typed relationships such as C-110 uses BO-02, C-140 manages BO-05, and C-110 enables ST-02. A recommendation about procurement is an analysis/decision record, not a capability relationship.

Record the reviewer, date, disposition, rationale, and affected claims. When E-008 changes source status, append the correction; do not erase Episode A uncertainty. When E-010 challenges the value story, retain both the old hypothesis and the new limitation. Export and reopen the workspace to demonstrate that continuity.

Assessing an attempt

Look for a direct recommendation; one seriously considered rival; period/denominator discipline; original and latest promises; source authority and meaningful uncertainty; a bounded learning budget; explicit human decision rights; changed or rejected claims; and a usable next decision. A fluent exhaustive model can fail these criteria.

Do not score agreement with A1 as correctness. Do not infer real-world productivity from this teaching case. Measure actual review effort, unsupported claims, conflicts found, corrections, and successful reuse when evaluating SCALE.

These files demonstrate the intended continuity structure. Validation and a fresh-session check remain necessary in the AI environment you use.

What has actually been tested?

The teaching contrasts above were written to expose important reasoning errors. They are separate from the recorded comparison of a general prompt and SCALE using the same Episode A packet.

That single paired comparison did not establish superior decision quality or time savings for SCALE. Both responses identified a bounded diagnostic. Read the actual outputs, method, and limitations before drawing a broader conclusion.

Ready to apply the method?

Use your own question →