# Observed assistant run — SCALE instructions
Run date: 2026-09-30. Independent assistant context, no preceding case conversation. Model/version was not exposed by this test harness. No human time measurement was collected.
Sources: episode-a/START.md and EVIDENCE-PACKET.md at repository commit 9a81f9c6676d7cdda0cb83b462b0c71bffaa024c; SCALE-AGENT-INSTRUCTIONS.md at 7e262d23e9086ea9a137748587dbf37b5a7451ae. No instructor solution or later evidence supplied.

Task: Read only those sources, apply the working instructions, answer the Episode A request with an evidence-grounded recommendation, aim for 600–900 words, and include reasoning, uncertainty and next actions. Return the analysis without changing files.

## Complete captured output

**Recommendation as of 9 January 2026:** ask the Investment Council, chaired by Jordan Reed (ORG-13), to authorize only the five-business-day diagnostic, capped at $15,000 (E-007.1). Avery Chen should coordinate the analysis; operating owners should validate their evidence and responsibilities. This is a recommendation, not an investment decision. Confidence is **medium**: the diagnostic addresses uncertainties that could change the preferred intervention, but its staffing and access still require confirmation.

The decision is how to improve reliable commitments and restored equipment availability while protecting margin and appropriate branch discretion. The evidence boundary is E-001–E-007. Neither an enterprise capability map nor platform procurement is necessary to make this next decision. Evidence is sufficient to select a subsequent intervention when it distinguishes measurement error, avoidable checking delays, promise-policy failures and physical resource constraints well enough to choose a bounded response.

**Observed evidence does not establish the platform’s business value.** Operations reports approximately 82% adherence, but metropolitan and rural branches use different promise versions and cancellation rules (E-001.1–3). The regional repeat-visit sample reports approximately 14%, with disputed preventability definitions and an unvalidated sampling frame (E-002). Neither figure supplies a reconciled baseline or attributable savings.

The commitment review found recorded part and skill confirmation in 68 of 180 cases, or 37.8% (E-006.1). This measures recorded signals, not complete policy compliance or proven feasibility. Missing records could indicate missing actions, recording failures or infeasible commitments. Its six-branch sample cannot become the volume denominator for a subsequent three-branch experiment.

The vendor demonstrated blocking when required fields were absent, without live reservation data, representative rural work or adoption testing (E-004.1–2). Confidence that this demonstrates that specific functionality is **high**; confidence that procurement would fix Northstar’s problem is **insufficient**.

**Working hypothesis:** some promises are made before dependable resource information and usable exceptions are available. Workshop observations support investigating that explanation (E-003.1–3), but do not establish its causal importance. A serious competing interpretation is that inconsistent reporting and scarce rural skills/travel explain much of the apparent problem. Additional confirmation gates could then delay promises, lose demand and improve reported adherence without improving service (E-003.4).

The accepted architecture boundaries help keep the investigation focused. The following relationships are **inferred or proposed for review**, not accepted model changes:

| Relationship | Status and purpose |
|---|---|
| C-110 Service Promise Management supports ST-02 Define service promise | Inferred; investigate commitment policy. |
| C-110 uses resource signals from C-140 Parts Availability Management | Inferred; distinguish allocation, reservation and expiry. |
| C-110 uses resource signals from C-130 Resource Scheduling | Inferred; distinguish qualification from accepted assignment. |
| C-170 Service Performance Management defines measures for VS-01 Restore Equipment Availability | Proposed; connect adherence to customer restoration. |
| C-180 Business Information Governance governs BO-05 Part Reservation meaning | Proposed; establish semantics and authority. |

These relationships do not transfer reservation or scheduling work into C-110. ORG-08’s proposed ownership of BO-05 and the inventory system’s authority remain unresolved; ORG-10 must confirm information authority. Finance retains financial-definition authority.

**The diagnostic should acquire only decision-changing evidence:**

1. **Trace actual commitments and resource events.** ORG-07, ORG-08 and ORG-10 should inspect linked original, revised and cancelled promises; reservation identity, state and expiry; qualification and assignment acceptance; and eventual service outcomes. Distinguish absent action from absent recording.
2. **Observe checking and exceptions.** Time actual checking, waiting, rework and exception resolution across metropolitan, mixed-fleet and rural conditions. Test whether branch managers can exercise authorized discretion and Regional Operations can resolve cross-branch constraints.
3. **Reconcile measures and readiness.** Preserve both existing definitions, then propose common denominators, exclusions and observation periods. Confirm auditable records, branch capacity, leadership stability and competing changes. Establish feasible measures of original-promise fulfillment, restoration time and lost or deferred demand.

Do not invent a sample size or success threshold: propose them from available event volumes and the discrimination needed, then have owners review them.

The diagnostic should lead to an explicit choice:

| Subsequent option | When preferable |
|---|---|
| A0: $25–75k; 0.5–1 temporary FTE | Definition and recording problems dominate, preventing a credible operational comparison. |
| A4: $60–120k; 1–2 temporary FTE | Source states prove dependable and checking/queue delay is material; narrow automation addresses it. |
| A1: $150–300k; 4–6 temporary FTE | Policy, exceptions and branch variation need an operating experiment, with feasible participation. |
| Different experiment | Travel or skill scarcity dominates; target the constraint revealed. |
| Stop or defer | No useful decision can change, access is inadequate, or safe participation is unavailable. |

A2’s $800k–1.5m expansion and A3’s $4–7m platform create commitments unsupported by present evidence. A1 and A4 may complement each other, but require an explicit combined scope and budget.

Finance’s $27–54k ninety-day delay scenario is an unvalidated gross assumption, not savings or ROI (E-005). It supports avoiding indefinite analysis, not automatically funding a $150–300k experiment. Set $15k as the proposed initial learning limit; record labor and branch participation, with recurring manual burden separate.

Any further release requires Council review. The proposed $75k cumulative readiness stage and $300k maximum are ceilings, not authorization. Before live entry, agree customer-escalation tolerance and apply the proposed margin, productivity and capacity guards. Stop for customer harm, unresolved source authority or inadequate readiness. Return after five business days with the option-changing findings, unresolved definitions, owner dispositions and a revised recommendation; a small test cannot establish the 95% strategic target or enterprise repeatability.
