Performance & Optimization - Testing & Continuous Improvement

Why Average Latency Misleads Backend Performance Newcomers

A backend developer moving into web performance will often start with the slowest API endpoint. That is a reasonable instinct and usually the wrong default. An endpoint can get faster while the page still feels slow, because rendering, JavaScript, and network scheduling may dominate the user’s wait. I would start with a real user action, then tune the part of its path that actually holds it up.

A fast API can still produce a slow page

Backend dashboards are persuasive because they expose timings you can act on: request duration, database spans, CPU time, and cache hit rate. None of those measures when a person can see or use the page, because the browser still has to download resources, parse code, lay out content, and respond to input. If you optimize an API before seeing that sequence, you are choosing a familiar bottleneck rather than a demonstrated one.

Google’s published “good” thresholds are a Largest Contentful Paint (LCP) of 2.5 seconds or less and an Interaction to Next Paint (INP) of 200 milliseconds or less, assessed at the 75th percentile. They are starting points for interpreting user experience, not targets you can satisfy by lowering server time alone. LCP may be late because the largest image is discovered after a JavaScript bundle runs; INP may be poor because a click waits behind main-thread work even though its API call is quick.

Consider an illustrative trace in which the API’s recorded p95 is 90 milliseconds while the page’s LCP is 3.1 seconds. Those figures are an example, not a claim about your application. If Chrome DevTools shows the LCP image starting late, cutting another 20 milliseconds from a SQL query cannot account for most of the visible delay. The useful question is no longer “Which endpoint is slowest?” but “What prevented this specific action from finishing sooner?”

You can read High Impact Performance Tuning for Modern Web Apps and still reject a server-first work queue, because a catalogue of effective changes cannot tell you which dependency is holding up your page. Begin with a named journey, such as opening a product page from a search result or saving an edited form. Record the URL, device class, network conditions, and the moment at which the user considers the action complete. Without that boundary, a faster server response is easy to celebrate even when the user sees no improvement.

The browser waterfall should decide whether backend work starts

For a first pass, use Chrome DevTools’ Performance and Network panels together. The Network waterfall tells you when a resource was requested and received; the Performance recording tells you what the browser did before it could paint or handle input. Lighthouse can suggest likely problems, but its lab run should not overrule field data because one scripted navigation cannot represent every device, connection, or returning visitor. The web-vitals library and the PerformanceObserver API can help collect browser measurements from actual sessions, provided you segment them rather than hiding divergent experiences in one average.

Separate the journey into waits you can assign to an owner: navigation and redirects; Time to First Byte (TTFB); resource discovery and transfer; JavaScript execution; rendering; and interaction processing. This is not a demand to optimize every stage. It is a way to avoid asking a database engineer to fix a late-loading hero image. If LCP waits on an image, check its request start, priority, format, and dimensions before touching a query. If INP is poor, inspect long tasks around the interaction before blaming the endpoint the click eventually calls.

A backend developer should still verify the server boundary, because a browser waterfall cannot explain what happened inside a request. The following Bash script uses curl to print TTFB and total transfer time for a URL supplied as its argument. Run it against a representative route, not only a cheap health check:

#!/usr/bin/env bash
set -euo pipefail
url="${1:?Usage: ./timings.sh URL}"
for i in {1..10}; do
  curl --fail --silent --show-error --output /dev/null \
    --write-out "run=$i ttfb=%{time_starttransfer}s total=%{time_total}s\n" \
    "$url"
done

The script takes 10 samples as a quick diagnostic, not as a statistically reliable percentile. It also measures from your machine, so it cannot stand in for mobile users on distant networks. Compare its result with the browser’s navigation timing: if both show a long wait for the first byte, server investigation is justified; if transfer ends early and paint happens much later, stay in the browser trace. Check redirects and CDN behavior as well, because a request served from an edge cache follows a different path from one reaching your application.

End-to-end traces beat a blanket caching sprint

Once the waterfall points at the server, the popular next move is to add caching. I would not make caching the first backend change for most new investigations, because a cache can conceal a slow path while adding invalidation rules, stale responses, and different behavior for warm and cold users. First establish whether the time is spent before your service, inside it, or after it. A Server-Timing response header can expose named server durations to browser tooling, while a W3C Trace Context traceparent header can connect a request to a distributed trace when propagation is configured.

OpenTelemetry instrumentation, a trace backend such as Jaeger, and Prometheus request-duration histograms answer related but different questions. A trace can show that one navigation invokes several sequential services; a histogram can show whether that pattern is common enough to matter. Do not compare a browser LCP percentile directly with a backend p95 and call the difference “frontend time,” because the two populations and clocks may differ. Instead, correlate a sampled navigation with its request traces, then check aggregated field measurements to see whether that navigation represents a recurring problem.

There is a concrete choice between response caching and critical-path simplification. Response caching wins when many requests repeat the same expensive, safely reusable result; its cost is cache storage, invalidation work, and the risk of serving data past its intended freshness. Critical-path simplification—removing a serial dependency or returning only data needed for the first render—wins when each request is distinct or when calls wait on one another; its cost is code and contract changes that need careful review. Neither option wins merely because it is familiar.

For cacheable public responses, inspect Cache-Control directives such as max-age and stale-while-revalidate alongside your CDN’s observed hit rate. Treat a provisional 60-second freshness lifetime as a value to tune against the data’s update requirements, not as a universal recommendation. For a serial call chain, inspect span start and end times before parallelizing anything, because apparent dependencies may enforce ordering or protect an overloaded downstream service. You can read Performance Tuning Tips for Faster Software Systems and still insist on path-specific evidence, because general tuning advice cannot identify which delay your users encounter.

One user journey is a better optimization target than a global score

A single site-wide performance score encourages work on whichever regression is easiest to display, because it compresses unlike journeys into one number. A landing page’s LCP, a search page’s filtering interaction, and a form’s save confirmation have different completion points. Choose one journey with meaningful traffic or a costly failure mode, and write down what “done” means before changing code. For a form, that might be visible confirmation after a successful save; for a landing page, it might be the main content appearing and becoming usable.

Give the journey a small measurement contract. Record field LCP or INP where relevant, the network request that gates completion, the corresponding server trace, and the user segment. A proposed 180 KB compressed JavaScript budget for the route is a tuning value, not a standard: it is useful only if reducing that payload shortens the observed wait on the devices you care about. Brotli compression, HTTP/2 or HTTP/3, and code splitting may affect delivery, but their presence alone does not prove the main thread can execute the downloaded code promptly.

Make one change at a time where practical, because simultaneous cache, query, and bundle changes make attribution difficult. Compare before and after for the same route and comparable users; keep an eye on error rate and correctness, because a page that paints quickly but fails to save is not an improvement. If a backend fix lowers request duration without moving the journey’s completion time, keep the server benefit if it matters for capacity, but stop describing it as a user-perceived speed win. That distinction makes the next optimization decision clearer.

Measure a journey before opening a tuning ticket

Tomorrow, pick one slow action reported by a user and capture a Chrome DevTools performance recording while completing it. Mark the moment the action feels finished, find the resource or task immediately before that moment, and follow its request into the backend only if the evidence points there. Write the bottleneck and its measurement on the ticket. That first trace is more useful than a backlog titled “make the API faster.”