The Curriculum / Reader / Organization-Wide Observability
LEVEL 3 · ADVANCED · COMPANY TRACK

Organization-Wide Observability

This page compiles 4 files from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-3-advanced/company/09-org-wide-observability/README.md

Organization-Wide Observability

Every AI team needs a shared truth about production behavior. Centralize schemas, lineage, access controls, and core metrics while preserving each product team’s domain dashboards and experiment velocity.

At company scale, AI is not a side project owned by the person who writes prompts. It is a production capability spanning product, domain operations, platform engineering, security, privacy, finance, and support. Belle Realty should make the operating decision visible: what is authorized, who is accountable, what evidence is required, and how the system is stopped when reality disagrees.

Operating model

Name one directly responsible individual for the outcome and one executive sponsor for the risk. Define the decision rights for data access, model changes, vendor changes, policy exceptions, and emergency shutdowns. Maintain a registry of deployed features with purpose, customer cohorts, model and retrieval versions, data sources, tool permissions, SLOs, evaluation evidence, and review date.

Control cadence

Review leading indicators weekly: reliability, safety interventions, access denials, cost, drift, user corrections, and unresolved incidents. Review material changes before launch and at a fixed expiry date after launch. The agenda should end in decisions, owners, and dates—not a dashboard tour.

Evidence standard

Require representative offline evaluation, staged release evidence, traceable telemetry, and a documented rollback path. Segment results by geography, property type, tenant context, language, and risk level. A global pass rate is insufficient for a system that operates differently for one high-risk cohort.

Practical scenario

Before allowing an assistant to send maintenance messages across a portfolio, prove authorization boundaries, approved language, escalation behavior, provider fallback, outage handling, audit retention, and per-property cost limits. Have operations rehearse the manual path and the stop path.

Decision test

The program is mature only when a new engineer can determine what is running, why it is permitted, how it is measured, and who can change or stop it without relying on tribal knowledge.

level-3-advanced/company/09-org-wide-observability/centralized-eval-registry.md

Centralized Evaluation Registry

Register every meaningful eval with purpose, owner, dataset provenance, version, metric definitions, thresholds, known blind spots, and linked decisions. Reuse cases and attacks instead of rebuilding a private test set per team.

At company scale, AI is not a side project owned by the person who writes prompts. It is a production capability spanning product, domain operations, platform engineering, security, privacy, finance, and support. Belle Realty should make the operating decision visible: what is authorized, who is accountable, what evidence is required, and how the system is stopped when reality disagrees.

Operating model

Name one directly responsible individual for the outcome and one executive sponsor for the risk. Define the decision rights for data access, model changes, vendor changes, policy exceptions, and emergency shutdowns. Maintain a registry of deployed features with purpose, customer cohorts, model and retrieval versions, data sources, tool permissions, SLOs, evaluation evidence, and review date.

Control cadence

Review leading indicators weekly: reliability, safety interventions, access denials, cost, drift, user corrections, and unresolved incidents. Review material changes before launch and at a fixed expiry date after launch. The agenda should end in decisions, owners, and dates—not a dashboard tour.

Evidence standard

Require representative offline evaluation, staged release evidence, traceable telemetry, and a documented rollback path. Segment results by geography, property type, tenant context, language, and risk level. A global pass rate is insufficient for a system that operates differently for one high-risk cohort.

Practical scenario

Before allowing an assistant to send maintenance messages across a portfolio, prove authorization boundaries, approved language, escalation behavior, provider fallback, outage handling, audit retention, and per-property cost limits. Have operations rehearse the manual path and the stop path.

Decision test

The program is mature only when a new engineer can determine what is running, why it is permitted, how it is measured, and who can change or stop it without relying on tribal knowledge.

level-3-advanced/company/09-org-wide-observability/executive-dashboards.md

Executive AI Dashboards

Show business value, reliability, safety, cost, exposure, and decisions needed. Include trends, targets, error-budget status, material incidents, and leading-risk indicators. Hide operational noise behind drill-downs.

At company scale, AI is not a side project owned by the person who writes prompts. It is a production capability spanning product, domain operations, platform engineering, security, privacy, finance, and support. Belle Realty should make the operating decision visible: what is authorized, who is accountable, what evidence is required, and how the system is stopped when reality disagrees.

Operating model

Name one directly responsible individual for the outcome and one executive sponsor for the risk. Define the decision rights for data access, model changes, vendor changes, policy exceptions, and emergency shutdowns. Maintain a registry of deployed features with purpose, customer cohorts, model and retrieval versions, data sources, tool permissions, SLOs, evaluation evidence, and review date.

Control cadence

Review leading indicators weekly: reliability, safety interventions, access denials, cost, drift, user corrections, and unresolved incidents. Review material changes before launch and at a fixed expiry date after launch. The agenda should end in decisions, owners, and dates—not a dashboard tour.

Evidence standard

Require representative offline evaluation, staged release evidence, traceable telemetry, and a documented rollback path. Segment results by geography, property type, tenant context, language, and risk level. A global pass rate is insufficient for a system that operates differently for one high-risk cohort.

Practical scenario

Before allowing an assistant to send maintenance messages across a portfolio, prove authorization boundaries, approved language, escalation behavior, provider fallback, outage handling, audit retention, and per-property cost limits. Have operations rehearse the manual path and the stop path.

Decision test

The program is mature only when a new engineer can determine what is running, why it is permitted, how it is measured, and who can change or stop it without relying on tribal knowledge.

level-3-advanced/company/09-org-wide-observability/shared-metrics-platform.md

Shared Metrics Platform

Publish standardized events for requests, traces, policy decisions, model usage, costs, quality labels, and outcomes. Enforce tenant-aware access and a canonical definition layer so executives do not compare incompatible numbers.

At company scale, AI is not a side project owned by the person who writes prompts. It is a production capability spanning product, domain operations, platform engineering, security, privacy, finance, and support. Belle Realty should make the operating decision visible: what is authorized, who is accountable, what evidence is required, and how the system is stopped when reality disagrees.

Operating model

Name one directly responsible individual for the outcome and one executive sponsor for the risk. Define the decision rights for data access, model changes, vendor changes, policy exceptions, and emergency shutdowns. Maintain a registry of deployed features with purpose, customer cohorts, model and retrieval versions, data sources, tool permissions, SLOs, evaluation evidence, and review date.

Control cadence

Review leading indicators weekly: reliability, safety interventions, access denials, cost, drift, user corrections, and unresolved incidents. Review material changes before launch and at a fixed expiry date after launch. The agenda should end in decisions, owners, and dates—not a dashboard tour.

Evidence standard

Require representative offline evaluation, staged release evidence, traceable telemetry, and a documented rollback path. Segment results by geography, property type, tenant context, language, and risk level. A global pass rate is insufficient for a system that operates differently for one high-risk cohort.

Practical scenario

Before allowing an assistant to send maintenance messages across a portfolio, prove authorization boundaries, approved language, escalation behavior, provider fallback, outage handling, audit retention, and per-property cost limits. Have operations rehearse the manual path and the stop path.

Decision test

The program is mature only when a new engineer can determine what is running, why it is permitted, how it is measured, and who can change or stop it without relying on tribal knowledge.

← Cost Governance at Scale The Continuous-Learning Organization →