# Observed assistant run — general analysis
Run date: 2026-09-30. Independent assistant context, no preceding case conversation. Model/version was not exposed by this test harness. No human time measurement was collected.
Sources: episode-a/START.md and EVIDENCE-PACKET.md at repository commit 9a81f9c6676d7cdda0cb83b462b0c71bffaa024c. No instructor solution or later evidence supplied.

Task: Read only those sources, answer the Episode A request with an evidence-grounded recommendation, aim for 600–900 words, and include reasoning, uncertainty and next actions. Return the analysis without changing files.

## Complete captured output

As of January 9, I recommend that the Investment Council consider **a five-business-day diagnostic capped at $15,000**, followed by an explicit decision on whether to stop, improve measurement, automate a narrow constraint, or test operating practices. This is the smallest evidenced commitment capable of changing the next decision (E-007.1). It does not authorize the broader experiment or platform. Avery should prepare and record the recommendation; the Council retains funding authority.

The proposed scheduling platform has **not been shown to fix Northstar’s service-promise problem**. Its demonstration established that it can block commitments when required fields are absent. It did not establish that populated fields are trustworthy, commitments are feasible, branches will adopt the controls, or customers regain usable equipment sooner (E-004).

**Observations.** Reported promise adherence is approximately 82%, but branches combine different rules for original versus revised promises and cancellations. This is not a comparable pilot baseline (E-001). Reported avoidable repeat visits are approximately 14%, with disputed classifications and an unvalidated sampling frame; the apparent gap to 8% cannot support attributable savings (E-002). Only 68 of 180 reviewed commitments had recorded prior part and skill confirmation. That 37.8% measures recorded checks, not complete policy compliance or proven infeasibility, and the six-branch sample cannot establish future three-branch volume (E-006).

**Hypotheses.** Premature promises and unreliable reservation meaning may cause failures: an allocation is not necessarily a physical reservation, and reservations can expire. However, a competing interpretation deserves equal investigation: much of the apparent failure may reflect recording differences, while actual disruption is driven by technician scarcity, travel, or slow exception authority. Stricter checks could delay commitments and lose customers without improving restoration. Workshop accounts support investigating these explanations; they do not rank their causal importance (E-003).

The relevant architecture is a bounded working model:

| Relationship | Status and implication |
|---|---|
| VS-01 concerns restored equipment availability; ST-02 defines a feasible commitment. | Accepted boundaries. Customer restoration must remain the outcome, alongside adherence. |
| C-110 uses signals from C-140 and C-130 to govern BO-02. | Inferred dependency requiring validation. C-110 does not reserve parts or schedule technicians. |
| BO-05 reservation states and BO-06 qualification/assignment checks inform promise feasibility. | Proposed validation relationship. Reservation authority remains unresolved; qualification and acceptance are separate. |
| C-180 supports common meanings; C-170 supports comparable measures and guards. | Proposed supporting relationships within accepted boundaries. ORG-10 confirms information authority; Finance retains financial authority. |
| I-01 connects these changes. | Proposed change boundary, not an approved solution or complete enterprise map. |

Operational accountability remains with branch, regional, field-service, and supply leaders.

The alternatives remain open:

| Option | Planning cost / peak temporary capacity | Decision implication |
|---|---|---|
| A0: reporting reconciliation | $25–75k / 0.5–1 FTE | Prefer if measurement uncertainty dominates and records cannot yet support an operating test. |
| A1: three-branch operating experiment | $150–300k / 4–6 FTE | Prefer if policy, exceptions, and branch variation require live learning and participation is safe. |
| A4: narrow feed and exception queue | $60–120k / 1–2 FTE | Prefer if reliable source states exist and manual checking is a material, removable delay. |
| A2: twelve-branch expansion | $800k–1.5m / 10–15 FTE | Defer: core practice and measurement remain unproven. |
| A3: enterprise platform | $4–7m / 20–30 FTE, 18–24 months | Defer procurement: functionality evidence does not justify the commitment. |

A1 and A4 may complement one another, but their costs and scopes should not be combined automatically.

The diagnostic should produce only the evidence needed to choose among these routes:

1. **Trace reservation and commitment events:** reconcile allocation, reservation, expiry, release, timestamps, and branch updates; distinguish missing actions from missing records. Supply Operations and Information Governance should resolve authority.
2. **Measure checking and exception work:** observe actual time and failures, including rural travel, skills, assignment acceptance, and approval delays. Determine whether automation addresses the bottleneck or merely reports it.
3. **Reconcile outcome definitions:** preserve original and revised commitments; resolve eligibility, cancellations, periods, repeat-visit coding, and links to customer disruption. Define restoration-time measurement before claiming improvement.
4. **Verify readiness and cost:** confirm branch leadership, competing change, auditable records, participation capacity, total labor effort, and recurring manual burden.

If scarce skills or travel dominate, design a narrower capacity or exception experiment. If source meaning remains unreliable, resolve that before live automation. Stop or redesign if no plausible finding would change the next decision.

Any later A1 proposal should request only the next justified tranche: up to $75,000 cumulative for readiness, and at most $300,000 cumulative only after operational entry. Proposed guards require agreement: no more than 3% unexplained productivity deterioration, six peak temporary FTE, Finance-governed margin review, and explicit customer-escalation tolerance. Unclear authority, unsafe capacity, or customer harm can stop entry or continuation regardless of analytical confidence (E-007).

The $27–54k ninety-day exposure is an unvalidated gross scenario excluding lost demand, not savings or ROI (E-005). It supports prompt learning, not automatic spending. A first release would test feasibility; it would establish neither enterprise repeatability nor achievement of the 95% strategic target.
