The Curriculum / Reader / Level 4 shared materials
LEVEL 4 · PROFESSIONAL

Level 4 shared materials

This page compiles 4 files from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-4-professional/shared/model-card-template.md

Model Card Template

Model identity

Intended use

Describe target users, supported tasks, operating languages, input limits, and expected environment. State prohibited or unsupported uses plainly.

Training and adaptation

Evaluation

Evaluation Dataset/slices Metric Result Limitations

Include comparison baseline, contamination controls, human-evaluation protocol, calibration, safety/refusal tests, and representative failures.

Risks and mitigations

Describe hallucination, bias, privacy, harmful-content, security, misuse, and distribution-shift risks. Pair every mitigation with its scope and remaining limitation.

Deployment notes

Specify supported quantizations, context limits, inference hardware, expected latency/cost, monitoring, change control, and incident contact.

Version history

Record changed data, weights, behavior, evaluations, and known regressions for every release.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-4-professional/shared/safety-case-template.md

Safety Case Template

A safety case is an evidence-backed argument that the residual risk of a specific system and use is acceptable. It is not a policy statement.

1. System and decision

2. Claim

State the claim narrowly: “For [use], under [conditions], the system can be operated with acceptable risk because [controls and evidence].”

3. Hazard register

For each hazard, record the harmed party, severity, likelihood, signals, prevention, mitigation, and residual risk.

Hazard Scenario Preventive controls Detection Response Residual risk

4. Evidence

5. Operating controls

Describe access control, human approvals, runtime policy, logging, monitoring, incident ownership, rollback, and user recourse.

6. Limits and review

List untested conditions, assumptions, expiry date, next review trigger, and the named authority who can suspend the system. Link raw evidence; do not summarize away uncertainty.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-4-professional/shared/system-card-template.md

System Card Template

A model card describes a model. A system card describes the complete deployed product: model, prompts, retrieval, tools, users, policies, interfaces, and operations.

System definition

Architecture

Describe input handling, identity/authorization, retrieval, model routing, tool calls, output validation, approval steps, logs, and data retention. Include a diagram link where available.

Behavior and boundaries

List supported tasks, refusal conditions, mandatory escalations, prohibited actions, localization requirements, and user disclosures. Explain what the system must not infer or decide.

Evidence

Claim Evaluation/method Result Owner Expiry

Include real-world monitoring and known failure examples, not only offline scores.

Risk operations

State red-team coverage, policy-engine controls, alert thresholds, incident playbook, rollback method, provider-outage plan, audit trail, and review cadence.

Change log

Every change to model, prompt, retrieval corpus, tools, permissions, or safety policy should have a linked evaluation and approval record.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-4-professional/shared/training-run-checklist.md

Training Run Checklist

Use this before any fine-tune, preference-optimization job, or material retraining. A run that cannot be explained, reproduced, evaluated, and rolled back is not ready.

Decision and scope

Data

Configuration and execution

Evaluation and release

Afterward

Write what changed, what did not, what surprised you, and which next experiment this result justifies. Negative results are assets when their conditions are preserved.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

← README Level 4 ELI10 companions →