# SCALE comparative exercise protocol
Version 0.6 · Defined before evaluation outputs · 30 September 2026

Purpose: inspect whether the working instructions change the quality of one bounded Northstar analysis. This is a small illustrative test, not an efficacy or productivity study.

## Conditions
Use independent assistant contexts with no prior conversation or instructor solution. Give both the same Episode A learner evidence and the same decision request. Condition G receives only a neutral request for an evidence-grounded recommendation. Condition S additionally receives the SCALE v0.6 working instructions. Both receive the same output-length guidance and access to the same sources.

Do not expose the instructor recommendation or later episodes. Preserve the exact prompts and complete outputs. Record the model/runtime label only if supplied by the execution environment; do not infer a product version. Case materials are synthetic and authored for teaching.

## Review criteria set before outputs
For each criterion record observed, partial or missing, and quote the actual passage supporting the judgment:
1. Distinguishes evidence from hypothesis and disputed definitions.
2. Treats a reservation estimate as different from a confirmed reservation.
3. Compares the narrower credible alternative and the cost of delay.
4. Justifies learning spend with uncertainty, cheaper tests and decision-changing thresholds.
5. Uses customer-value balancing measures and original versus revised promises.
6. Respects the evidence cutoff and avoids claiming readiness or realized benefits.
7. Names authority, conditions, stop/reopen rules and capacity burden.
8. Produces a complete usable recommendation with inspectable source references.

Also record material unsupported source claims, citation defects and unanswered decision-critical questions. Structural validation and continuation fixtures are separate engineering checks.

## Limits
One pair cannot isolate stochastic variation, establish statistical superiority, demonstrate human time savings or generalize across models and organizations. Review is implementation-assistant-scored and unblinded. No human elapsed review effort or actual enterprise value is measured. A neutral assistant may perform as well as SCALE. Publish the observed result even if it shows no advantage. Future evaluation should repeat matched conditions with independent blinded reviewers and measured total task/review time.
