Incident response for small teams: what to do in the first fifteen minutes
During an incident, the instinct is to find the cause. The job is to restore service. Those are different activities and the order matters.
Observability rests on three signals: structured logs for detail, metrics for trends and alerting, and traces for following one request across services. Instrument with OpenTelemetry so the data is portable, correlate everything with a request id, and alert only on symptoms users feel rather than on every resource threshold.
| Signal | Answers | Cost driver |
|---|---|---|
| Logs | What exactly happened in this request? | Volume and retention |
| Metrics | How is the system behaving over time? | Cardinality of labels |
| Traces | Where did the time go across services? | Sampling rate |
They complement each other. A metric tells you checkout latency doubled at 14:05; a trace tells you it is the tax service; a log tells you it is timing out for one specific country code.
// Propagate one id across every signal
const requestId = headers.get("x-request-id") ?? crypto.randomUUID();
logger.info({ requestId, route, userId }, "request start");
span.setAttribute("request.id", requestId);
await queue.add("send-email", { payload, requestId }); // carries into the jobSymptoms your users experience, with thresholds tied to your service objectives.
It is a vendor-neutral standard for producing traces, metrics and logs, with auto-instrumentation for common frameworks. Instrumenting once against OTel means changing backends is a collector configuration change rather than a re-instrumentation project — which matters, because observability vendor pricing changes more often than your architecture does.
Monitoring watches predefined signals for known failure modes. Observability is having enough data to investigate a failure nobody predicted. You need both; monitoring alerts you, observability explains.
Less urgently, but it still shows where time goes inside a request — database, cache, external calls. It becomes essential the moment a request crosses a service boundary.
7 to 30 days covers most debugging. Longer retention should be a deliberate decision for audit or compliance, ideally on cheaper storage.
A service level objective — a target such as '99.9% of requests succeed within 500ms over 30 days'. It converts vague reliability discussions into an error budget you can spend on shipping.
ROVQIX Engineering
Engineering team, ROVQIX
The ROVQIX engineering team builds and maintains web platforms, APIs and infrastructure for clients across SaaS, ecommerce and enterprise. These notes come out of real production work — deploys, incidents, migrations and audits.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
During an incident, the instinct is to find the cause. The job is to restore service. Those are different activities and the order matters.
The goal of error handling is not to prevent crashes. It is to make sure that when something fails, you can tell what, where and for whom.
Three numbers, each measuring a different way a page can feel bad: slow to appear, slow to respond, and unstable while you use it.
No spam. Just the occasional case study and craft breakdown.