GROUND
Problem
What becomes confusing, fragile, or impossible without understanding observability? This lesson answers that through explanation, a worked example, two runnable exercises, and a reference solution. No teacher-supplied worksheet is required.
Distributed architectures address concrete capacity, reliability, and coordination limits—but often arrive before those limits do.
LEARN
Concept explanation
Observability belongs to “Distributed systems without cargo cult”. observability is inspected by introducing one realistic failure and proving root cause before repair.
For observability, trace concrete input, state transition, output, and failure through an architecture decision tied to explicit requirements, estimates, state ownership, failure modes, and measured limits.
Start with requirements and estimates, choose the smallest architecture, locate its first bottleneck, then evolve one measured constraint at a time. Apply that model to supplied normal, boundary, and failure cases; each case below names its input and expected evidence.
Evidence produced by the observability experiment: output, state, trace, bytes, timing, or diagnostics.
Condition that must remain true while inputs or implementation of observability change.
Point where observability crosses ownership, representation, time, process, network, or trust.
Example bank
Compare normal, boundary, failure, and cross-layer cases. Predict each observation before revealing the explanation.
SETUPCapture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time.
OBSERVEEvidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
WHY IT MATTERSThis isolates the normal contract of observability; preserve its raw evidence as the control for every later comparison.
SETUPRecord observability behavior at: Model partition and replica lag in user-visible terms.
OBSERVERecord what remains invariant and the first representation, owner, size, or timing value that changes in Architecture notebook · load tests · traces · Git history.
WHY IT MATTERSA boundary example is useful only when one named dimension changes and everything else stays comparable.
SETUPDiagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.
OBSERVECapture the first divergence from the baseline, including exact input, diagnostic, state, and recovery result. Expected recovery: Evidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
WHY IT MATTERSThe diagnostic is part of the interface. Repair the proven cause, not the most visible symptom.
SETUPTrace observability one layer below its usual abstraction through an architecture decision tied to explicit requirements, estimates, state ownership, failure modes, and measured limits.
OBSERVETrace requests, state ownership, capacity, queues, caches, replication, failure domains, recovery, operational load, and cost.
WHY IT MATTERSThe lower layer is earned when it explains evidence the current layer cannot. Otherwise keep observability at the simpler boundary.
SEE
Worked example
Start from supplied design.md. Focus: Capture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time.
- Run: Fill design.md, then verify every component traces to one stated requirement or measured constraint.
- Save baseline evidence. Trace requests, state ownership, capacity, queues, caches, replication, failure domains, recovery, operational load, and cost.
- Boundary case: Record observability behavior at: Model partition and replica lag in user-visible terms.
- Failure case: Diagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.
RESULT
Evidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax. Starter-level baseline: Document names one action, target, non-goal, numeric estimate, state owner, request path, bottleneck, and recovery procedure.
START HERE
Starter material
PREREQUISITESA plain-text editor, Git, and tools named by the weekly slice. Start from a blank directory.
ONE-TIME SETUPmkdir reforging-capstone && cd reforging-capstone && git init
Create design.md, paste this exact content, then run the command below.
# observability
## Requirements
- One primary user action:
- One reliability target:
- One explicit non-goal:
## Estimate
- Requests/second:
- Stored bytes/day:
- Peak concurrent work:
## Smallest design
- State owner:
- Request path:
- First bottleneck:
- Recovery procedure:
Fill design.md, then verify every component traces to one stated requirement or measured constraint.STOP / CLEANUPStop any process started by your slice with Ctrl+C; run git status before leaving.
DO WITH GUIDANCE
Guided exercise
Diagnose a controlled failure: observability
- Normal case: Capture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time.
- Write predicted evidence from this named case before running starter.
- Break one assumption, capture evidence, then repair only proven cause.
- Run exact normal case. Save commands, inputs, outputs, and diagnostics in notebook.
- Explain changed evidence using lesson mental model in no more than five sentences.
Concrete guided solution
- Copy the supplied design.md unchanged and run: Fill design.md, then verify every component traces to one stated requirement or measured constraint.
- Write this prediction before inspecting output: Evidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
- Perform only the named normal case: Capture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time.
- Save the raw output, then annotate input → transition → evidence. Use Architecture notebook · load tests · traces · Git history to confirm the transition rather than inferring it.
- Compare prediction with evidence; if they differ, keep both and write the rule that explains the difference. Reference baseline: Document names one action, target, non-goal, numeric estimate, state owner, request path, bottleneck, and recovery procedure.
DO ALONE
Independent exercise
Prove root cause: observability
- Create second case from blank file: Record observability behavior at: Model partition and replica lag in user-visible terms.
- Then create controlled failure: Diagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.
- Use Architecture notebook · load tests · traces · Git history to prove behavior, then repair controlled failure.
- Compare result against supplied acceptance checks and reference approach before marking complete.
Concrete independent solution
- Duplicate the starter into a clean comparison case; change only this boundary: Record observability behavior at: Model partition and replica lag in user-visible terms.
- Save its evidence beside the baseline and identify the first changed value. Trace requests, state ownership, capacity, queues, caches, replication, failure domains, recovery, operational load, and cost.
- Create the exact controlled failure: Diagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.
- Save healthy trace, output, metadata, or state snapshot for observability: Design chat/search from single node, then add one measured distributed constraint at a time.
- Write one cause hypothesis and evidence that would falsify it.
- Trigger exact failure without adding other changes: Lose leader or duplicate delivery and specify recovery/invariant.
- Diff healthy and failed evidence; repair first divergent boundary only.
- Run boundary plus baseline again and confirm: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
- Rerun baseline, boundary, and repaired failure together. Accept only if all reproduce: Evidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
COMPARE
Expected result
- Evidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
- Document names one action, target, non-goal, numeric estimate, state owner, request path, bottleneck, and recovery procedure.
- Controlled observability failure produces captured evidence; repair restores stated invariant without hiding error.
PROVE
Acceptance checks
Lesson is complete only when every check is true. Each check is stored locally and travels with your JSON backup.
0/5 complete · saved on this device
UNSTICK
Hints
Reveal hints
- Start with supplied normal case exactly as written: Capture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time.
- For boundary case, change only named dimension: Record observability behavior at: Model partition and replica lag in user-visible terms.
- If result is confusing, diff raw inputs and evidence before editing implementation.
- If tool shows nothing useful, move observation one boundary lower: representation, runtime, OS, or network.
VERIFY
Solution
Attempt both exercises before opening reference approach.
Reveal reference solution
- Run unmodified starter and preserve baseline evidence: Document names one action, target, non-goal, numeric estimate, state owner, request path, bottleneck, and recovery procedure.
- Save healthy trace, output, metadata, or state snapshot for observability: Design chat/search from single node, then add one measured distributed constraint at a time.
- Write one cause hypothesis and evidence that would falsify it.
- Trigger exact failure without adding other changes: Lose leader or duplicate delivery and specify recovery/invariant.
- Diff healthy and failed evidence; repair first divergent boundary only.
- Run boundary plus baseline again and confirm: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
PREDICT · INSPECT · BREAK · DEBUG · MEASURE
Interrogate reality
Prediction: write expected output, state transition, ordering, and failure evidence before running either exercise.
Inspection: Trace requests, state ownership, queues, caches, replication lag, failure domains, cost, and recovery procedures.
Measurement: Estimate traffic, storage, latency, throughput, availability, consistency windows, operational load, and cost.
Capture raw evidence before explaining.
Change one assumption and force controlled failure.
Find cause with Architecture notebook · load tests · traces · Git history before editing fix.
MASTERY + FRONTIER + BOUNDARY
Own the knowledge
Explain observability at beginner, intermediate, and senior depth.
Recreate smallest useful example from blank file without notes or AI.
Schedule recall for day 1, 7, 30, and 90.
Creative frontier lab
Try first without opening the solutions. The constraints invite invention; the reference gives one concrete direction, never the only valid answer.
Re-solve observability by removing the most convenient abstraction. delete one service and recover the requirement with the smallest capable layer.
CONSTRAINTKeep the same inputs, observable result, and failure evidence; change the means, not the contract.
ORIGINAL IDEATurn subtraction into a design tool: the missing abstraction should reveal which responsibility it used to hide.
Reveal frontier solution
- Freeze the contract as three fixtures: Capture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time. / Record observability behavior at: Model partition and replica lag in user-visible terms. / Diagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.
- List every convenience used by the starter; remove the highest-level one while preserving Fill design.md, then verify every component traces to one stated requirement or measured constraint..
- Implement the smallest replacement using delete one service and recover the requirement with the smallest capable layer.
- Run all fixtures and compare raw evidence. Keep the simpler version unless the removed abstraction has a demonstrated benefit.
Build an explanation artifact for observability: trace each component to a requirement, estimate, owner, failure domain, and observable signal.
CONSTRAINTA peer must be able to locate the first divergence without reading implementation code.
ORIGINAL IDEATreat the explanation itself as a product: make invisible transitions visible, replayable, and diffable.
Reveal frontier solution
- Create one row or timestamped event for each transition in: Capture healthy observability baseline from: Design chat/search from single node, then add one measured distributed constraint at a time.
- For every row record input, representation, owner, operation, output, and tool evidence from Architecture notebook · load tests · traces · Git history.
- Replay Record observability behavior at: Model partition and replica lag in user-visible terms.; highlight only changed rows.
- Replay Diagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.; stop at the first divergent row and attach its recovery action.
Combine the boundary and failure into a new user-visible scenario for observability. turn a failure drill into a user-visible recovery story with explicit invariants.
CONSTRAINTDo not merely add more input. Invent a recovery interaction, alternate representation, or self-checking behavior.
ORIGINAL IDEAMake the system teach its own limits: the artifact should expose the invariant and offer a safe next action when it breaks.
Reveal frontier solution
- Combine these two pressures without changing them: Record observability behavior at: Model partition and replica lag in user-visible terms. AND Diagnose this exact observability failure before editing: Lose leader or duplicate delivery and specify recovery/invariant.
- Name the invariant that must survive and the user-visible evidence when it cannot: Evidence identifies first divergence, repair changes only proven cause, and rerun restores: Diagram names state owner, consistency window, failure domain, trace path, and coordination tax.
- Implement this original direction: turn a failure drill into a user-visible recovery story with explicit invariants.
- Demonstrate baseline, combined failure, recovery, then baseline again; save the sequence as a regression fixture.
Capability frontier
Push observability until another layer becomes justified. Record one robust technique, one contextual trade-off, and one labeled hack or historical curiosity.
CORE · PRACTICAL · CONTEXTUAL · HACK · FRAGILE · HISTORICAL · GOLF
Boundary
No diagram removes physics, coordination, product ambiguity, or operational responsibility.
If this vanished tomorrow…
Collapse toward one process, one machine, files, SQLite, explicit protocols, and manual recovery until the true minimum appears.
Why next layer is earned
The next layer is another year of rebuilding, publishing, source reading, and measured production work.