Specifies thresholds, routing, dashboards, diagnostic steps, escalation, remediation, and review for health checks. This reference fixes the adoption boundary, failure behavior, and observable evidence while leaving environment-specific limits configurable.
Package status: reference context ready for human review. The contract and test scenarios are complete, but no claim is made that an adopting implementation has passed them.
Decision
For Health checks, adopt a versioned telemetry contract tied to service objectives and runbooks. Operate the capability with named signals, bounded capacity, owned alerts, a diagnostic runbook, and tested recovery evidence. This package is a reference contract: the adopting team must choose numeric limits, providers, jurisdictions, and operational owners from evidence in its own environment.
Scope
- The Health checks actors, inputs, outputs, states, policy or schema versions, and externally visible outcomes described by this operations block.
- Primary and alternate interfaces, background work, caches, integrations, support paths, and evidence that can exercise the same boundary.
- Adoption-specific configuration, rollout, recovery, and verification responsibilities needed to apply the reference safely.
Outside this block
- Choosing a universal vendor, framework, jurisdiction, numeric threshold, retention period, or service-level objective for every adopter.
- Claiming that packaged scenarios have run against a downstream implementation or that this reference grants legal, security, or accessibility certification.
Contract
- Define each signal, unit, aggregation, attributes, owner, expected cardinality, retention, and decision it supports.
- Propagate correlation across request, queue, job, and dependency boundaries without placing secrets or unrestricted personal data in attributes.
- Alert on user-impacting symptoms or exhausted safety margins, with a diagnostic path and a named response owner.
- Treat sampled traces and client reports as evidence; do not let unauthenticated telemetry authorize production mutation or rollback by itself.
- Define normal, degraded, exhausted, and unavailable states with decision metrics and explicit operator actions.
- Preserve accepted-work identity during failover, replay, rollback, and repair so recovery cannot duplicate or silently lose effects.
- Separate startup, liveness, readiness, and deeper diagnostic checks and define the exact local invariant, timeout, frequency, failure threshold, and orchestrator action for each.
- Keep routine probes cheap and free of secrets, avoid coupling liveness to every downstream dependency, and test cold start, drain, partial dependency loss, overload, recovery, and version mismatch.
Implementation guidance
- Practice the runbook in a representative environment and record recovery time, data loss boundary, and unresolved assumptions.
- Model Health checks inputs, outputs, actors, states, invariants, side effects, and evidence before selecting framework or vendor details.
- Store the applicable Health checks contract or policy version with material state so migrations, replay, and support decisions remain attributable.
- Introduce the path behind controlled rollout, compare expected and observed outcomes, and keep a tested recovery route until adoption evidence is complete.
Failure handling and safeguards
- Missing telemetry is itself detectable, while the monitored service remains bounded and does not depend on observability export success.
- Invalid or contradictory Health checks state is rejected or quarantined; the implementation never guesses a value that broadens authority or duplicates an effect.
- Retries are bounded and use durable identity; after exhaustion, work reaches an inspectable terminal state with an accountable owner.
Verification and operations
- Inject saturation, dependency outage, stale configuration, and telemetry loss and confirm the declared degraded behavior.
- Monitor Health checks success, denial, validation failure, dependency failure, retry exhaustion, and recovery by contract version without sensitive payload dimensions.
- Re-run paired positive and negative fixtures after policy, schema, dependency, migration, or boundary changes and preserve the resulting evidence.
Adoption assumptions
- Names and values in the example are a concrete fixture, not universal defaults; adopters replace them through documented evidence and ownership.
- The adopting system can provide authenticated identity, durable operation or record versions, bounded telemetry, and a controlled path for change.
The executable-looking examples in this package are fixtures and acceptance contracts. Run the collection validator to check structure and metadata, then translate and execute the scenarios in the target repository before recording implementation evidence.