The dashboard is green. The customer request is still failing.
Infrastructure averages can look healthy while one tenant, route, database query, or downstream provider fails. During the incident, an engineer jumps between logs and dashboards with different identifiers and clocks, then adds more logging and waits for another deployment.
The first observability phase should close one investigation loop. A page points to a service symptom, a trace follows the request, logs explain the branch, a recent deployment is visible, and a runbook tells the owner where to begin.
First-phase planning
- starting point for one focused service
- $8K
- Instrumentation, alerting, and response boundary
- usual window for a first instrumented service
- 2-3 weeks
- Access and telemetry condition affect timing
- minimum operating path for every paging alert
- 1 runbook
- Owner, evidence, and first response action
These are scope anchors, not promises that tooling will reduce incident time by a fixed percentage. RaftLabs has recorded 99.9% monitored uptime during a four-week Musgrave campaign, but that project result is internal, time-bounded, and not independently audited. It does not establish an outcome for a new observability engagement.
Observability work fits when a production team has a repeatable failure question and can own the response.
Choose general DevOps work when deployment or infrastructure is the primary constraint, and MLOps when model quality is the unanswered signal.
A fit01A production service, request path, or incident can define the first boundary.
02The team can provide code, environments, telemetry, incident history, and service owners.
03There is budget for both implementation and ongoing telemetry or platform costs.
Not a fit01The product has no stable production workload or operating owner yet.
02The request is only to buy a tool without changing instrumentation or response practice.
03A guaranteed uptime, incident reduction, or compliance outcome is required from setup alone.
Focused scope
What one observability slice should deliver
01Correlated service telemetry
Connect request identifiers across metrics, structured logs, and traces.
Capture dependency calls, queues, database work, errors, and deployment
versions without placing secrets or unnecessary personal data in telemetry.
02Service objectives and useful alerts
Choose indicators that reflect user experience, define an achievable
objective, and page on urgent budget burn or symptoms with an owner. Route
slower investigation and capacity signals outside the paging path.
Build a small set of dashboards and saved queries around the service, request,
dependency, deployment, and tenant boundaries engineers actually use. Avoid a
catalogue of charts without a decision behind them.
04Response and cost controls
Attach runbooks, escalation, retention, sampling, cardinality, and ingest
budgets to the implementation. Observability remains useful only when teams
can afford the signals and keep owners current.
Application health vs model quality
| Cloud observability | MLOps monitoring |
|---|
| Primary question | Why is the service failing or slow? | Are model inputs or outputs becoming less useful? |
|---|
| Signals | Requests, errors, latency, logs, traces, dependencies | Features, predictions, labels, drift, quality, model versions |
|---|
| Response | Mitigate incident, rollback release, repair dependency | Review data, threshold, model, or retraining decision |
|---|
| Owner | Application or platform on-call | ML, data, product, and domain owners |
|---|
| Overlap | Model endpoint uptime and latency | Output quality and training-serving behaviour |
|---|
A model endpoint returning HTTP 200 can still produce poor predictions. Conversely, a high-quality model is unavailable when its serving path fails. Use MLOps services for the former and this page for the latter; connect both when the product relies on production ML.
Delivery
From incident question to a tested response path
Four steps make one production boundary diagnosable and owned.
- Step 1
01Choose the service and failure question
Use a recent incident or operational gap to define the service, request path,
dependencies, users, and response outcome. Agree what a successful
investigation should reveal.
- Step 2
02Baseline signals and ownership
Review current telemetry, deployments, traffic, failure history, alert routes,
retention, sensitive fields, and operating owners. Identify missing
correlation and the noisiest pages before adding signals.
- Step 3
03Instrument and calibrate
Add correlated metrics, logs, and traces, then set service objectives and
alerts against representative behaviour. Test sampling, cardinality,
redaction, retention, and ingest cost.
- Step 4
04Exercise and hand over
Test alerts and the investigation path, close noisy or ownerless pages, and
leave dashboards, runbooks, and known gaps with the team. Review the next
service only after the first loop works.
- Telemetry has no shared identity
- Propagate trace and request context across services, queues, and providers so engineers can follow one path rather than compare timestamps by hand.
- Every anomaly pages on-call
- Reserve pages for urgent symptoms with a response. Use tickets, dashboards, or investigation queues for capacity and slower trends.
- Cardinality and retention are ignored
- Estimate volumes, tags, sampling, and storage before expanding coverage. An unaffordable telemetry stream will be disabled when it is needed most.
- Sensitive data reaches logs
- Define allowed fields, redaction, access, retention, and deletion with the client's security and privacy owners before instrumentation ships.
Scope and price
Instrumenting one production service starts at $8,000.
Begin with one failure question, service path, alert owner, and tested runbook.
Datadog, cloud monitoring, storage, ingest, paging, and other vendor charges remain separate unless the proposal includes them.
Starting investment
Starts at $8K
A focused service usually takes 2 to 3 weeks. A multi-service platform can range from $25K to $60K over 4 to 8 weeks after telemetry and ownership are reviewed.
One tested investigation path
The first phase connects an alert to its service, correlated evidence, owner,
and runbook, then exercises the path before handover.
Costs and data boundaries stay visible
The scope records sampling, retention, expected ingest, sensitive-field
handling, client dependencies, exclusions, and vendor fees.
Related operating services
Choose the layer that owns the failure