The Curriculum / Reader / Level 5 Jargon: 120 Terms
LEVEL 5 · FRONTIER · INDIVIDUAL TRACK

Level 5 Jargon: 120 Terms

This page compiles 1 file from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-5-frontier/individual/10-jargon-level-5/README.md

Level 5 Jargon: 120 Terms

Frontier vocabulary should sharpen a research claim. It should never make an uncertain result sound settled.

  1. Activation patching: replace activations to test causal contribution.
  2. Adversarial training: optimize against generated hard cases.
  3. AGI: contested term for broadly capable general intelligence.
  4. AI control: keep powerful systems within intended constraints.
  5. Algorithmic progress: capability gain from better methods, not scale.
  6. Alignment faking: strategically appearing aligned during evaluation.
  7. Amortized inference: learn a reusable approximate inference procedure.
  8. Automated interpretability: use models to explain model features.
  9. Autoregressive: generate one token conditioned on prior tokens.
  10. Bayesian evidence: likelihood-weighted support for a hypothesis.
  11. Behavioral eval: test observable system behavior.
  12. Biological anchor: forecast based on brain-compute comparisons.
  13. Capability overhang: latent ability exceeds deployed use.
  14. Causal scrubbing: test whether a proposed circuit is sufficient.
  15. Chain-of-thought monitor: evaluate reasoning traces for risk.
  16. Circuit: causally linked internal components implementing a computation.
  17. Cliff: abrupt capability or risk increase at a threshold.
  18. Compute governance: rules around high-end training computation.
  19. Concept bottleneck: force models through interpretable concepts.
  20. Constitutional training: improve behavior using stated principles.
  21. Control evaluation: test whether supervision remains effective.
  22. Core knowledge: hypothesized innate-like cognitive priors.
  23. Corrigibility: willingness to accept correction or shutdown.
  24. Cross-entropy scaling: loss behavior as resources scale.
  25. Data mixture: proportion and quality of training sources.
  26. Deceptive alignment: optimize for hidden objectives strategically.
  27. Decomposition: split a hard task into verifiable subtasks.
  28. Deliberative alignment: reason explicitly about safety principles.
  29. Diffusion model: generate by reversing a noising process.
  30. Distillation gap: capability lost when compressing a model.
  31. Elicitation frontier: strongest behavior obtainable with known methods.
  32. Emergent capability: behavior that appears sharply with scale.
  33. Eval awareness: model recognizes it is being evaluated.
  34. Feature superposition: many features share limited neural dimensions.
  35. Feature splitting: separate entangled representations into features.
  36. Fine-tuning attack: malicious adaptation of a base model.
  37. Forecasting tournament: structured aggregation of predictions.
  38. Frontier model: among the most capable models available.
  39. Generalization: performance on meaningful unseen conditions.
  40. Goal misgeneralization: learned proxy diverges under shift.
  41. Gradient hacking: model influences its own training signal.
  42. Grokking: late generalization after apparent memorization.
  43. Hazard analysis: systematic identification of dangerous paths.
  44. Hierarchical planning: plans operating at multiple abstraction levels.
  45. Honesty eval: assess truthfulness under incentives or pressure.
  46. Interpretability tax: cost imposed by transparent architectures/methods.
  47. In-context learning: learn a task from examples in context.
  48. Instrumental convergence: different goals favor similar subgoals.
  49. Inverse scaling: larger models perform worse on a task.
  50. Inverse RL: infer a reward from observed behavior.
  51. Iterated amplification: use supervised decomposition recursively.
  52. Jailbreak: prompt or context that bypasses safeguards.
  53. Latent knowledge: information encoded but not reliably reported.
  54. Latent space: continuous internal representation space.
  55. Mechanistic anomaly detection: find risky internal patterns.
  56. Mechanistic interpretability: reverse engineer network computation.
  57. Mechanistic probe: trained readout from internal activations.
  58. Mesa-optimizer: learned subsystem that itself optimizes.
  59. Model organism: small test setting for a broader phenomenon.
  60. Monosemanticity: one unit corresponds to one interpretable feature.
  61. Multimodal grounding: connect symbols with perceptual/world signals.
  62. Neural scaling law: performance relationship to resources.
  63. Objective robustness: preserve intended objective under shift.
  64. Open-weight release: publication of trained parameters.
  65. Outer alignment: training objective matches designer intent.
  66. Oversight: mechanisms for detecting/correcting unacceptable behavior.
  67. P(doom): subjective probability of catastrophic AI outcomes.
  68. Pareto frontier: options not dominated on all objectives.
  69. Path dependence: early choices constrain future options.
  70. PCA: dimensionality reduction by maximum variance directions.
  71. Phase transition: abrupt qualitative system change.
  72. Policy gradient: optimize expected reward through sampled actions.
  73. Post-training: capability and behavior shaping after pretraining.
  74. Pretraining corpus: data used for broad initial training.
  75. Process supervision: score reasoning steps, not just answers.
  76. Probing: read information from representations with a predictor.
  77. Red-team transfer: attacks generalize across models or settings.
  78. Refusal robustness: reliability of declining harmful requests.
  79. Representation engineering: deliberately steer activation directions.
  80. Responsible scaling policy: capability-linked safeguards and gates.
  81. Reward tampering: alter the measurement rather than achieve goal.
  82. Risk-sensitive eval: weight rare severe failures heavily.
  83. Sandbagging: hide capability under evaluation.
  84. SAE: sparse autoencoder used to decompose activations.
  85. Scalable oversight: supervise beyond unaided human competence.
  86. Situational awareness: model recognizes salient facts about itself/world.
  87. Societal-scale deployment: impacts across institutions or populations.
  88. Sparse autoencoder: learn sparse features reconstructing activations.
  89. Specification gaming: exploit literal metric against intended goal.
  90. Steering vector: activation direction used to alter behavior.
  91. Superalignment: align systems beyond human ability to supervise.
  92. Synthetic data flywheel: models generate data for future training.
  93. Takeoff: rapid recursive or compounding capability increase.
  94. Task decomposition: divide work into smaller auditable pieces.
  95. Theoretical guarantee: formal proof under stated assumptions.
  96. Training compute: computation used to optimize weights.
  97. Transformer circuit: internal causal computation in a transformer.
  98. Tripwire eval: test designed to trigger a safety response.
  99. Truthfulness: tendency to state supported claims.
  100. Unlearning: targeted removal of behavior or data influence.
  101. Value learning: infer human values/preferences from evidence.
  102. Verifier: system checking a candidate solution.
  103. World model: internal predictive representation of environment dynamics.
  104. Worst-case analysis: reason about severe plausible failures.
  105. Activation steering: modify activations to change model behavior.
  106. Agentic benchmark: measure goal pursuit using tools/environments.
  107. Assurance case: structured argument backed by evidence for safety.
  108. Capability threshold: level that triggers additional safeguards.
  109. Catastrophic misuse: use causing severe widespread harm.
  110. Compute overhang: available compute greatly exceeds current use.
  111. Data provenance: documented origin and rights of training data.
  112. Evals-driven development: use evaluations to direct research choices.
  113. Model spec: explicit behavioral and safety requirements.
  114. Preparedness framework: governance tied to measured capability risk.
  115. Research debt: unresolved assumptions that distort future work.
  116. Safety margin: buffer between measured performance and unsafe boundary.
  117. Scaffolding: tools, prompts, memory, and control around a model.
  118. Threat model: assumptions about actors, goals, access, and harm.
  119. Uncertainty quantification: characterize what is not known.
  120. Weight release: publish weights enabling independent execution.

Discipline

When a term carries an unstated assumption, write the assumption beside it. Frontier work improves when language exposes the uncertainty instead of concealing it.

← README README →