The Curriculum / Reader / Observability at Depth
LEVEL 3 · ADVANCED · INDIVIDUAL TRACK

Observability at Depth

This page compiles 5 files from the repository, verbatim, in reading order. The living version: this folder on GitHub.

level-3-advanced/individual/06-observability-at-depth/README.md

Observability at Depth

Logs tell you something happened; traces explain the path. For AI, retain the lineage from request through retrieval, model calls, policy decisions, tools, response, and later user outcome.

Treat this as an operating document, not a reading assignment. For Belle Realty or a property-management assistant, define the unit of work, owner, allowed failure modes, and the customer-visible consequence before choosing a model or framework. A system that gives plausible answers but cannot be measured, contained, or recovered is not production-ready.

Design stance

Start with a narrow contract. State what enters the boundary, what the component may read or change, what it must return, and when it must stop. Put the contract in version control alongside representative tenant, listing, maintenance, and leasing cases. Prefer a deterministic rule when one exists; use model judgment only where language or ambiguity genuinely adds value.

Operating loop

Instrument every request with a correlation ID, model and prompt version, retrieval version, tool calls, latency, token usage, policy decisions, and final outcome. Review a small, stratified sample of real traffic weekly. Track failures by property, workflow, language, channel, and tenant impact; aggregate averages hide the exact bad experience that produces churn.

Controls

Set an explicit threshold that changes behavior: block, require approval, degrade to search-only, route to a cheaper model, or page an owner. Test the threshold with synthetic failures and recent production examples. Do not make a dashboard metric a promise unless a named person can act on it within a defined window.

Production exercise

Apply this to a maintenance triage request: a resident reports a gas smell at 2 a.m. Document the safe path, the tool permissions, the human escalation, the message shown to the resident, and the evidence retained for review. Ship only after the behavior is repeatable in a rehearsal.

Exit criteria

The implementation has a written contract, measurable outcomes, a failure path, an accountable owner, and a rollback or containment action. If any is missing, it is still a prototype.

level-3-advanced/individual/06-observability-at-depth/custom-metrics.md

Custom AI Metrics

Track domain outcomes: verified showing scheduled, maintenance emergency correctly escalated, correct tenant record selected, answer accepted, and harmful action prevented. Token count is a cost input, not a quality metric.

Treat this as an operating document, not a reading assignment. For Belle Realty or a property-management assistant, define the unit of work, owner, allowed failure modes, and the customer-visible consequence before choosing a model or framework. A system that gives plausible answers but cannot be measured, contained, or recovered is not production-ready.

Design stance

Start with a narrow contract. State what enters the boundary, what the component may read or change, what it must return, and when it must stop. Put the contract in version control alongside representative tenant, listing, maintenance, and leasing cases. Prefer a deterministic rule when one exists; use model judgment only where language or ambiguity genuinely adds value.

Operating loop

Instrument every request with a correlation ID, model and prompt version, retrieval version, tool calls, latency, token usage, policy decisions, and final outcome. Review a small, stratified sample of real traffic weekly. Track failures by property, workflow, language, channel, and tenant impact; aggregate averages hide the exact bad experience that produces churn.

Controls

Set an explicit threshold that changes behavior: block, require approval, degrade to search-only, route to a cheaper model, or page an owner. Test the threshold with synthetic failures and recent production examples. Do not make a dashboard metric a promise unless a named person can act on it within a defined window.

Production exercise

Apply this to a maintenance triage request: a resident reports a gas smell at 2 a.m. Document the safe path, the tool permissions, the human escalation, the message shown to the resident, and the evidence retained for review. Ship only after the behavior is repeatable in a rehearsal.

Exit criteria

The implementation has a written contract, measurable outcomes, a failure path, an accountable owner, and a rollback or containment action. If any is missing, it is still a prototype.

level-3-advanced/individual/06-observability-at-depth/langfuse-vs-langsmith-vs-helicone.md

Langfuse vs LangSmith vs Helicone

Choose by workflow, not branding: compare self-hosting and data control, trace and eval primitives, provider coverage, prompt management, cost visibility, access controls, and exportability. Keep a vendor-neutral event schema.

Treat this as an operating document, not a reading assignment. For Belle Realty or a property-management assistant, define the unit of work, owner, allowed failure modes, and the customer-visible consequence before choosing a model or framework. A system that gives plausible answers but cannot be measured, contained, or recovered is not production-ready.

Design stance

Start with a narrow contract. State what enters the boundary, what the component may read or change, what it must return, and when it must stop. Put the contract in version control alongside representative tenant, listing, maintenance, and leasing cases. Prefer a deterministic rule when one exists; use model judgment only where language or ambiguity genuinely adds value.

Operating loop

Instrument every request with a correlation ID, model and prompt version, retrieval version, tool calls, latency, token usage, policy decisions, and final outcome. Review a small, stratified sample of real traffic weekly. Track failures by property, workflow, language, channel, and tenant impact; aggregate averages hide the exact bad experience that produces churn.

Controls

Set an explicit threshold that changes behavior: block, require approval, degrade to search-only, route to a cheaper model, or page an owner. Test the threshold with synthetic failures and recent production examples. Do not make a dashboard metric a promise unless a named person can act on it within a defined window.

Production exercise

Apply this to a maintenance triage request: a resident reports a gas smell at 2 a.m. Document the safe path, the tool permissions, the human escalation, the message shown to the resident, and the evidence retained for review. Ship only after the behavior is repeatable in a rehearsal.

Exit criteria

The implementation has a written contract, measurable outcomes, a failure path, an accountable owner, and a rollback or containment action. If any is missing, it is still a prototype.

level-3-advanced/individual/06-observability-at-depth/slo-dashboards.md

SLO Dashboards

Dashboards need current status, error-budget burn, cohort breakdowns, deploy annotations, and direct links to traces. Do not put dozens of unlabeled charts in front of an on-call engineer.

Treat this as an operating document, not a reading assignment. For Belle Realty or a property-management assistant, define the unit of work, owner, allowed failure modes, and the customer-visible consequence before choosing a model or framework. A system that gives plausible answers but cannot be measured, contained, or recovered is not production-ready.

Design stance

Start with a narrow contract. State what enters the boundary, what the component may read or change, what it must return, and when it must stop. Put the contract in version control alongside representative tenant, listing, maintenance, and leasing cases. Prefer a deterministic rule when one exists; use model judgment only where language or ambiguity genuinely adds value.

Operating loop

Instrument every request with a correlation ID, model and prompt version, retrieval version, tool calls, latency, token usage, policy decisions, and final outcome. Review a small, stratified sample of real traffic weekly. Track failures by property, workflow, language, channel, and tenant impact; aggregate averages hide the exact bad experience that produces churn.

Controls

Set an explicit threshold that changes behavior: block, require approval, degrade to search-only, route to a cheaper model, or page an owner. Test the threshold with synthetic failures and recent production examples. Do not make a dashboard metric a promise unless a named person can act on it within a defined window.

Production exercise

Apply this to a maintenance triage request: a resident reports a gas smell at 2 a.m. Document the safe path, the tool permissions, the human escalation, the message shown to the resident, and the evidence retained for review. Ship only after the behavior is repeatable in a rehearsal.

Exit criteria

The implementation has a written contract, measurable outcomes, a failure path, an accountable owner, and a rollback or containment action. If any is missing, it is still a prototype.

level-3-advanced/individual/06-observability-at-depth/tracing-and-spans.md

Tracing and Spans

Use one trace per customer interaction and spans for retrieval, reranking, model inference, tool calls, guardrails, and writes. Propagate IDs across queues so delayed maintenance actions still join the original request.

Treat this as an operating document, not a reading assignment. For Belle Realty or a property-management assistant, define the unit of work, owner, allowed failure modes, and the customer-visible consequence before choosing a model or framework. A system that gives plausible answers but cannot be measured, contained, or recovered is not production-ready.

Design stance

Start with a narrow contract. State what enters the boundary, what the component may read or change, what it must return, and when it must stop. Put the contract in version control alongside representative tenant, listing, maintenance, and leasing cases. Prefer a deterministic rule when one exists; use model judgment only where language or ambiguity genuinely adds value.

Operating loop

Instrument every request with a correlation ID, model and prompt version, retrieval version, tool calls, latency, token usage, policy decisions, and final outcome. Review a small, stratified sample of real traffic weekly. Track failures by property, workflow, language, channel, and tenant impact; aggregate averages hide the exact bad experience that produces churn.

Controls

Set an explicit threshold that changes behavior: block, require approval, degrade to search-only, route to a cheaper model, or page an owner. Test the threshold with synthetic failures and recent production examples. Do not make a dashboard metric a promise unless a named person can act on it within a defined window.

Production exercise

Apply this to a maintenance triage request: a resident reports a gas smell at 2 a.m. Document the safe path, the tool permissions, the human escalation, the message shown to the resident, and the evidence retained for review. Ship only after the behavior is repeatable in a rehearsal.

Exit criteria

The implementation has a written contract, measurable outcomes, a failure path, an accountable owner, and a rollback or containment action. If any is missing, it is still a prototype.

← Adversarial Hardening Cost at Scale →