The Curriculum / Reader / Level 5 ELI10 companions
LEVEL 5 · FRONTIER

Level 5 ELI10 companions

This page compiles 6 files from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-5-frontier/eli10/what-is-a-frontier-model-eli10.md

What Is a Frontier Model?

A frontier model is like the fastest race car in the world right now: one of the most capable AI systems available. It may be unusually strong at writing, coding, planning, science, images, tool use, or other difficult tasks.

Being at the frontier does not mean it is perfect or wise. The fastest race car needs better brakes, trained drivers, careful rules, and a safe track. Powerful AI needs testing, access controls, monitoring, and people who can stop it when something goes wrong.

Frontier models are often trained with huge amounts of data and computing power. They can unlock useful work, such as helping researchers explore ideas or helping businesses automate complex tasks. They can also make mistakes at a larger scale or give more people access to powerful capabilities.

That is why responsible teams focus on measurements, not hype. What tasks can the model do? Under what conditions? How reliable is it? Can it use tools? What dangerous behavior has been tested? What happens if someone tries to misuse it?

The name “frontier” describes a moving edge. A model can be impressive today and ordinary next year. What matters is whether its benefits are real and its risks are handled before it is put on the road.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/eli10/what-is-a-research-lab-eli10.md

What Is a Research Lab?

A research lab is a team whose main job is to discover reliable new knowledge. It is not just a room full of smart people or a place that builds flashy demos. A good lab asks hard questions, designs fair tests, keeps careful records, and shares enough detail that others can check its work.

Imagine a workshop for finding out how the world works. Researchers start with an idea, make a prediction, build an experiment, and see whether the result agrees. If it does not, that is still useful: it tells them the idea needs to change.

An AI research lab might study better learning methods, how models reason, how to make systems safer, or how to understand what happens inside neural networks. It needs computers and data, but also good evaluation, honest writing, and a culture where people can say “the result did not work.”

The best labs combine ambition with discipline. They pursue important questions, but they do not announce more certainty than the evidence supports. They also consider who could be helped or harmed by what they discover.

At its best, a lab leaves behind more than products: methods, measurements, ideas, and people that make the whole field better.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/eli10/what-is-agi-eli10.md

What Is AGI?

AGI stands for artificial general intelligence. People use it to mean an AI that can handle many different kinds of thinking problems, learn quickly, and transfer what it knows from one area to another—more like a broadly capable person than a tool built for one narrow job.

There is no single agreed finish line. Does an AI need to match a person at every task? Be able to learn new jobs from a few examples? Work independently for long periods? Have common sense? Different people answer differently, which is why “we have AGI” is usually more of a claim than a measurement.

The important question is not the label. It is what the system can actually do, how reliably it does it, what tools it can access, how much human supervision it needs, and what harms become possible if it fails or is misused.

Think of it like calling a car “very fast.” That is less useful than knowing its top speed, brakes, crash tests, driver assistance, and where it is allowed to drive. As AI gets more capable, concrete evaluations and safeguards matter more than big labels.

AGI is a compass for a possible future, not a product requirement. Treat predictions about it with curiosity, evidence, and humility.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/eli10/what-is-mechanistic-interpretability-eli10.md

What Is Mechanistic Interpretability?

Mechanistic interpretability is like giving an AI a CAT scan. Instead of only checking what the AI says, researchers look inside its network to see which internal parts are active and how information moves through them.

Imagine a very complicated machine that can answer questions but has no instruction manual. You can press buttons and watch what happens, but that only tells you its outside behavior. A CAT scan helps you look inside without taking the whole machine apart. Researchers use special tools to inspect the AI’s internal numbers, patterns, and connections.

The goal is to find explanations that are more than a pretty picture. For example: “This group of connections notices a name,” or “changing this internal signal makes the model stop doing a certain kind of math.” To prove an explanation, researchers try changing a part and checking whether the predicted behavior changes too.

This work is hard. AI brains are not organized like human brains, and the same idea can be spread across many tiny pieces. A colorful chart does not automatically explain anything.

If it works well, mechanistic interpretability could help us find dangerous shortcuts, hidden capabilities, or reasons for mistakes before they hurt someone. Today, it is a promising scientific tool—not a finished safety system.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/eli10/what-is-pretraining-eli10.md

What Is Pretraining?

Pretraining is like teaching a kid to read by giving them a giant library—the whole internet, books, code, articles, and other text—then repeatedly asking, “What word probably comes next?” The child does not get a list of facts to memorize. They learn patterns from seeing enormous numbers of examples.

An AI model does something similar. It reads pieces of text broken into tokens and practices predicting the next token. After enough practice, it becomes good at grammar, facts, styles, reasoning patterns, and many kinds of tasks. That broad first education is called pretraining.

Pretraining does not make the AI a trustworthy expert. A child who has read every book can still misunderstand a question, repeat a mistake from a book, or confidently guess. After pretraining, people usually give the model extra teaching: examples of how to follow instructions, feedback on helpful answers, safety rules, and tools for checking current information.

The quality of the library matters. It must be collected legally and responsibly, cleaned, balanced, and checked for private information or repeated material. Bad or biased material can teach bad patterns at huge scale.

Pretraining creates general ability. The rest of the system determines whether that ability is useful, safe, and accountable in the real world.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

level-5-frontier/eli10/what-is-responsible-scaling-eli10.md

What Is Responsible Scaling?

Responsible scaling means putting seatbelts on before trying to drive at top speed. When an AI team makes a model more powerful, it should also increase the testing, security, oversight, and rules around that model.

Imagine a race car company. Before making the engine faster, it checks the brakes, tires, steering, and safety gear. It decides which speeds are safe on which tracks and stops racing if something important fails. AI teams can use the same idea.

They first measure what a model can do and what could go wrong. If it crosses a capability threshold—perhaps it becomes much better at cyber work, long independent tasks, or helping create dangerous things—the team adds stronger controls. Those might include more red-team testing, limited access, human approval, independent review, or delaying a release.

Responsible scaling does not mean never building better AI. It means refusing to treat capability and safety as separate races. The faster the system, the more evidence you need that people can understand, supervise, and contain its risks.

The hard part is honesty. Teams must be willing to slow down when an evaluation is inconclusive, a safeguard fails, or the potential harm is bigger than their ability to manage it. That is what makes the word “responsible” meaningful.

This lesson belongs in a practitioner’s operating system, not a collection of facts to recite. The point is to make a better decision under uncertainty: define the claim, identify the evidence that could change it, name the failure mode, and record the consequence of being wrong. Read it with a live initiative in mind—an internal workflow, customer-facing product, training run, or research bet—and turn the ideas into an explicit test.

Start from the outcome rather than the technology. Specify the user or stakeholder, the task boundary, the data and permissions involved, the success measure, and the unacceptable result. Establish a baseline before changing anything. Then make the smallest reversible move that can distinguish competing explanations. A plausible demo is evidence of possibility, not evidence of reliability, value, or safety.

Keep an evidence log. Separate observations from interpretations, measured performance from anecdotes, and known risks from assumptions. Review representative failures by hand; aggregate metrics can hide the one pattern that matters. For high-impact work, assign a clear owner, predefine an escalation path, and decide what will cause a pause or rollback. Do not outsource accountability to a model, vendor, benchmark, or committee.

The professional standard is legibility. Another capable person should be able to understand why this approach was chosen, rerun the evaluation, find its limits, and improve it without guessing. Build reusable artifacts—datasets, decision records, checklists, incident notes, and release criteria—so each project leaves the next one stronger.

Working exercise

Write a one-page decision memo for a current initiative. State the hypothesis, baseline, evaluation, threshold, owner, risks, and next action. If any of these cannot be stated plainly, the work is not ready to scale.

← Level 5 shared materials Contributing, and the license →