The Curriculum / Reader / Level 5 shared materials
LEVEL 5 · FRONTIER

Level 5 shared materials

This page compiles 3 files from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-5-frontier/shared/frontier-lab-charter-template.md

Frontier Lab Charter Template

Mission

State the capability or scientific aim and the public-benefit obligation in the same paragraph. A frontier lab without a concrete social contract will optimize only for momentum.

Scope and red lines

Define research domains, excluded work, capability thresholds requiring review, and situations in which the lab will pause, delay, or restrict release.

Governance

Research norms

Commit to reproducibility, candid limitation reporting, strong baselines, responsible disclosure, authorship standards, and preservation of negative results. Explain how researchers can challenge senior views without career penalty.

Operations

Specify compute controls, access management, dataset governance, evaluation gates, red-teaming, incident handling, vendor dependencies, and audit retention.

Public accountability

Publish a regular, concrete account of major capabilities, safeguards, evaluations, incidents where appropriate, and unresolved limitations. Do not claim certainty the evidence does not support.

Review

Set a review cadence, named signatories, amendment rules, and explicit triggers for re-evaluation: capability jumps, new misuse evidence, deployment expansion, or material regulatory change.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/shared/paper-writing-checklist.md

Paper-Writing Checklist

Before drafting

Argument

Integrity

Safety and release

Final review

Ask a skeptical reader to reproduce the central figure and restate the claim. If either task fails, revise before submission.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/shared/research-agenda-template.md

Research Agenda Template

North-star question

Write a question important enough to matter if solved and narrow enough to generate experiments within one quarter.

Why now?

Describe the new capability, observation, data access, theory, or institutional condition that makes this question timely. Do not confuse novelty with importance.

Map of beliefs

Hypothesis Supporting evidence Disconfirming evidence Next discriminating experiment

Research program

For each workstream, define the mechanism claim, minimal testbed, baseline, evaluation, budget, owner, and stop condition. Include one line for why the result would generalize—or why it should not.

Safety and social stakes

Name misuse possibilities, dual-use concerns, affected groups, external reviewers, and release constraints. Safety is part of the research design, not an appendix after a positive result.

Portfolio

Allocate time across exploitation (reliable progress), exploration (high variance), and infrastructure (reproducibility, data, evaluations). Revisit the allocation when evidence changes, not simply on a calendar.

Decision log

Link all launched, killed, and paused projects. A durable agenda preserves the reasons, not just the headlines.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

← README Level 5 ELI10 companions →