Try Northstar · Episode A of three
What should Northstar fund now?
Northstar wants more dependable customer promises. A scheduling platform has been demonstrated, but the operating problem and the right investment remain open questions.
Use your AI to produce a concise recommendation. Review its reasoning, then compare it with the reference answer.
Northstar v1.2 is a fictional teaching case. All organizations, source records, figures, and events are authored for the exercise.
- What to fund
- Ready to begin?
- What the results justify
Start with two files and one prompt.
- Download these two files.1. SCALE instructionsSCALE-AGENT-INSTRUCTIONS.md · How the assistant should work2. Episode A evidenceEVIDENCE-PACKET.md · Everything needed at this decision
Use only the initial investment evidence. The packet contains no reference answer or later results. Prefer one ZIP? Download the learner packet.
- Attach both files to a new AI conversation.
Use an AI that accepts Markdown attachments. Ask it to confirm that it can read both files. Keep the reference answers and later episodes out of this conversation.
- Copy this prompt and send it to your AI.
Episode A exercise prompt
Use the attached SCALE-AGENT-INSTRUCTIONS.md as the working method. EVIDENCE-PACKET.md contains fictional Northstar evidence available by January 9, 2026. Use only that evidence. Help me recommend the next bounded commitment in response to Northstar's service-promise problem. I am advising the Investment Council, which retains the authority to decide. Produce a concise decision brief, about one to two pages, that: - Gives a direct recommendation with specific source locators for material claims. - Compares at least one credible rival and explains its tradeoffs and opportunity cost. - Distinguishes observations, hypotheses, assumptions, and recommendations. - Identifies the important uncertainty, cheapest useful next evidence, cost and capacity limits, stop conditions, and what would change the recommendation. - Names the decision owner and any conditions requiring that owner's action. Use existing architecture only where it helps explain the decision. Preserve missing or conflicting information; do not invent facts, causal estimates, financial benefits, or approvals. Finish by identifying the conclusion most vulnerable to challenge and the evidence that could change it. I will review the draft. Keep my review separate from a fictional Council decision. A reviewed brief and a short correction note are sufficient; no JSON export or complete architecture model is required.
After your AI responds
Review the recommendation before comparing answers.
- Open the sources behind the important claims. Does the evidence support them?
- Challenge the strongest alternative, missing information, and anything that would change the recommendation.
- Correct the draft and keep your brief with its uncertainty, conditions, and named decision maker.
If you want to continue later, ask the AI for a short continuation note preserving your corrections, rejected claims, unresolved questions, and next step. Save it with your brief and source files.
Have a recommendation you can explain?
Compare the Episode A answer →The reference is one reasoned response, not the only acceptable conclusion. A different recommendation should explain its evidence and tradeoffs.
Evidence available at this decision
Read Episode A online.
The two-file setup’s evidence download combines everything needed for this episode. The reader below separates the current packet from any earlier context.
Northstar: evidence available before the investment decision
Learning edition: 1.2 · Decision date: 2026-01-09 · Classification: public, entirely fictional. Boundary: E-001–E-007 only. This packet contains no subsequent decision or result.
These are newly authored fictional source extracts for a teaching exercise. They extend Northstar v1.1 with additional uncertainty, alternatives, and financial assumptions; they are not recovered underlying records from v1.1 or observations from a real company. Quoted speakers and reviewers are fictional. Source authority establishes who supplied a claim, not that the claim is correct.
The request
Northstar Equipment Services has $780 million in annual revenue, 48 branches in twelve states, and 3,200 employees. It maintains and rents commercial equipment. Customers need equipment available when their own work requires it.
Priya Shah, EVP Service Operations (ORG-09), asks:
Build us a capability map and tell the Council whether the proposed scheduling platform fixes our service-promise problem.
Jordan Reed chairs the Investment Council (ORG-13), which controls funding. Avery Chen is the business architect (ORG-14); Avery can frame, analyze, recommend, and record, but cannot approve investment. Morgan Ellis, CFO (ORG-12), governs financial definitions.
Strategy records identify O-01: improve reliable service promises with appropriate branch discretion; O-02: protect margin; O-03: reduce customer disruption and repeat work. The executive target is 95% promise adherence by 2027. A possible first release would test feasibility, not establish the strategic target or enterprise repeatability.
Existing architecture fragments
The following identifiers are established; their relevance and relationships still need review. Do not invent a complete enterprise map.
| ID | Accepted name and boundary |
|---|---|
| VS-01 | Restore Equipment Availability; value belongs to the customer whose equipment becomes usable |
| ST-02 | Define service promise; move from accepted need to a understood, feasible commitment |
| C-110 | Service Promise Management; govern the commitment, using resource signals without reserving parts or scheduling technicians itself |
| C-130 | Resource Scheduling; match authorized work to qualified people, time, geography, and capacity |
| C-140 | Parts Availability Management; determine, reserve, position, consume, release, and expire part availability |
| C-170 | Service Performance Management; define and explain measures and guardrails, without acquiring Finance's authority |
| C-180 | Business Information Governance; govern business meaning, ownership, quality, and semantic change |
| BO-02 | Customer Promise; a commitment, not every ETA estimate; preserve original and revised versions |
| BO-05 | Part Reservation; proposed owner ORG-08 VP Supply Operations; inventory record is a candidate source, not confirmed authority |
| BO-06 | Technician Assignment; skill qualification and assignment acceptance are separate checks |
| I-01 | Connected Service Promise; proposed change boundary, not an approved solution |
ORG-02 branch managers control local capacity and authorized exceptions; ORG-05 Regional Operations resolves cross-branch constraints; ORG-07 Field Service owns field performance; ORG-08 Supply Operations owns parts performance; ORG-10 Information Governance confirms information authority. Architecture stewardship does not transfer operating accountability.
E-001
Source: Operations dashboard extract and definition notes, 2026-01-05; owner ORG-09; twelve-month reporting across all 48 branches. Directionally useful, not a reconciled pilot baseline.
| Source fragment | Recorded observation |
|---|---|
| E-001.1 enterprise dashboard | Reported M-01 promise adherence: approximately 82%. Local denominators were combined without a common cancellation rule. |
| E-001.2 metropolitan definition | Count completed work against the latest accepted appointment; remove customer cancellations. |
| E-001.3 rural definition | Count work against the original promised window unless an operational manager approves a revision; cancellations remain until reviewed. |
| E-001.4 customer desk note | Missed or changed promises generate complaints, but service-recovery cases are not consistently linked to the original commitment. |
Analyst warning: these definitions may produce different rates for identical service. Reconcile original versus latest promise, eligible population, revisions, cancellations, and reporting period before treating 82% as comparable to a pilot result. No customer-restoration-time baseline has been approved.
E-002
Source: Regional quality sample, 2026-01-06; owner ORG-07; repeat-visit review.
E-002.1: The sample reports approximately 14% avoidable repeat visits (M-02), against an 8% strategic aspiration. E-002.2: Two regions dispute whether customer-requested additional work, diagnostic return visits, and unavailable parts count as preventable. E-002.3: The sampling frame and event-level coding require validation. The apparent gap cannot yet support an attributable savings calculation.
E-003
Source: Five branch workshops and dispatcher observation notes, 2026-01-07; collected by Avery Chen; convergent observations, not a controlled causal study.
| Locator | Fictional source extract |
|---|---|
| E-003.1 dispatcher, metro | “I promise a date while the customer is on the line. The part screen says allocated; I call later to find out whether it is physically held.” |
| E-003.2 supply supervisor | “An allocation is a planning quantity. A reservation can expire. A screen refresh is not a guarantee that the part is still there.” |
| E-003.3 rural manager | “The dates are difficult because the qualified technician may be three hours away. Approval for an exception sometimes takes longer than making another plan.” |
| E-003.4 commercial manager | “If we wait for every confirmation before quoting a date, some customers will leave. Later promises may improve the dashboard while making our service worse.” |
| E-003.5 technology lead | “The existing system exposes a timestamp and reservation identifier. I have not established whether all branches update them consistently.” |
Competing explanations include premature commitment, unreliable reservation meaning, constrained travel/skills, unusable exception authority, and measurement artifacts. Neither these accounts nor a capability diagram establishes their relative causal contribution.
E-004
Source: Vendor demonstration and internal integration note, 2026-01-08; technology owner ORG-11.
E-004.1: In a controlled demonstration the proposed scheduling product blocked commitments when required part and skill fields were absent. E-004.2: No live Northstar reservation feed, representative rural work, adoption test, or service-value experiment was used. E-004.3: The vendor proposes $4–7 million over 18–24 months, subject to scope and commercial agreement. This is not an accepted quote. E-004.4: An internal team proposes a narrower reservation/skill-status feed and dispatcher exception queue using existing systems. Its $60–120k and one-to-two temporary FTE planning range assumes usable reservation identifiers and source-state semantics. It would not solve technician scarcity, travel, or disputed promise policy.
The narrow option is credible enough to assess. It is not proven feasible by the availability of an API.
E-005
Source: Finance planning memorandum, 2026-01-08; Morgan Ellis, ORG-12.
E-005.1: Governed enterprise gross service margin (M-03) is 31%; branch and work-mix variation are material. A proposed guardrail is at least 31% cumulative at investment reviews. E-005.2: The case has no approved causal estimate of avoidable loss. For planning only, Finance brackets recoverable service credits/overtime associated with the candidate problem at $9–18k per month across the intended test boundary. The assumption is unvalidated, gross, and excludes lost demand; do not count it as an achieved benefit or an ROI. E-005.3: A 90-day delay therefore exposes a scenario of $27–54k in those costs if nothing else changes. Some delay may be necessary to prevent a larger wrong investment; the range does not settle the decision. E-005.4: The $150–300k broad operating-test range must include internal labor, external work, training, branch participation, reconciliation, and measurement. Report recurring manual burden separately from one-time change cost. Peak FTE is a capacity constraint, not total labor effort. E-005.5: Before material commitment, identify which uncertain fact would change the next decision, what cheaper work can resolve first, and the maximum amount worth spending to learn it. No $300k authorization should be treated as an instruction to spend $300k.
E-006
Source: Two-week commitment review, 2026-01-08; provisional M-04 sample.
E-006.1: 68 of 180 reviewed commitments had recorded part and skill confirmation before commitment: 37.8%, conventionally reported as 38%. E-006.2: The sample spans six candidate branches and broad work types. It is not the later three-branch eligible cohort and cannot be annualized into pilot volume. E-006.3: This checks part and skill signals, not every policy obligation. A commitment can pass those checks and still lack required customer communication, geography validation, or a proper approval. E-006.4: The review does not establish whether missing evidence means a missing action, a recording failure, or a genuinely infeasible commitment.
E-007
Source: Branch capacity and options worksheet, 2026-01-09; Regional Operations with Finance and Delivery; planning evidence.
Candidate test conditions: one metropolitan branch, one mixed-fleet branch with part constraints, and one rural branch with long travel and scarce skill substitution. Select deliberately for variation, not statistical representation. Verify leadership stability, safe participation capacity, auditable records, and competing change before activation.
| Option | Scope | Planning cost / capacity | What it might establish | Material weakness |
|---|---|---|---|---|
| A0 | Reconcile measures and improve reporting only | $25–75k; 0.5–1 temporary FTE | Baseline comparability and problem extent | Does not directly change promise practices |
| A1 | Three-branch operating experiment using existing systems | $150–300k; 4–6 temporary FTE including measured branch participation | Policy/exception feasibility, manual burden, differentiated branch constraints | Requires scarce capacity; may not distinguish each cause |
| A2 | Immediate twelve-branch operating expansion | $800k–1.5m; 10–15 temporary FTE | Broader variation sooner | Commits before the core practice and measures are proven |
| A3 | Procure and deploy enterprise platform | $4–7m; 20–30 temporary FTE over 18–24 months | Technology fit and operating results after major commitment | High path dependence; present evidence proves functionality only |
| A4 | Targeted reservation/skill feed plus exception queue | $60–120k; 1–2 temporary FTE | Whether a narrow automation removes a material delay | Depends on unresolved source meaning; limited policy/capacity coverage |
A1 and A4 may be complements or alternatives. There is no requirement to recommend A1.
E-007.1: A five-business-day diagnostic is possible within $15k, charged within any eventual A1 cap: reconcile a sample of reservation events; time actual checking; inspect original/revised/cancelled commitments; test exception authority; verify available branch capacity. E-007.2: A prospective staged A1 envelope would release up to $15k for that diagnostic, up to $75k cumulative for policy/source/readiness work, and at most $300k cumulative only after an operational-entry decision. These are proposed controls, not recorded approval. E-007.3: Potential guards include M-06 no more than 3% unexplained productivity deterioration; M-08 at most $300k and six peak temporary FTE; customer-escalation tolerance agreed before release; original-promise fulfillment and time-to-restored-availability monitored alongside M-01. E-007.4: Operational readiness, customer harm, and unclear source authority can stop a test even after analysis is sufficient for an investment decision.
Your task
Recommend the smallest defensible next commitment, including a competing interpretation you take seriously. Identify what would make A0, A4, a different experiment, or stopping preferable. State whether your architecture relationships are observed, inferred, proposed, or accepted. Give the Council a usable decision brief and preserve unresolved definitions.
Do not manufacture a probability of success, monetary value of information, or causation. If the proposed learning cannot change a decision, redesign it before requesting the full experiment.