Grace
Full-stack engineer — distributed systems, observability, ML infrastructure

demos / ranger

Find the service that actually broke

A span's duration includes the time its children were running, so when something deep in a call tree slows down, every ancestor slows by exactly the same amount. Sort by what got slower and the whole request path is at the top of the list, all of it correct, none of it useful. Self time — the work a span did itself — is what separates the service that broke from the services waiting on it.

Incident
adManualGc
Operations ranked
87
Localized
ad · GetAds
Runner-up score
246 vs 1.25
Live report — click a trace to change the timeline open full width →

Real traces, not a simulation: two 300-second windows captured from the OpenTelemetry Demo at commit 8c47d47, with the demo's own adManualGc feature flag providing ground truth. The page is self-contained — every trace is embedded, nothing is fetched, and it opens from file:// on a laptop with no network, which is the state most people are in during an incident.

What to look at

Open the first trace in the table. The ancestors — load-generator, the proxy, the frontend — are drawn in grey, because their own work did not change; they are waiting. The red bar is ad · oteldemo.AdService/GetAds, whose self time went from 9.5ms to 2.36s. That is the entire claim, and it is visible in one screen.

The ranking on the left is scored by effect size — each operation's self-time shift as a share of its own baseline. The cause scores +24606%. The next operation down scores +125%, and it is a load-generator operation: the nearest competitor in this window is the test harness rather than any service.

Ranking ninety-odd operations and reporting the largest score is a maximum over ninety candidates, so the score alone does not say whether it is surprising. Shuffling the two windows' labels within each operation and re-ranking, 300 times, says what the best score reaches by chance here: 1.23. The observed score is 246. p=0.003, and the same seed gives the same answer every time.

This is one case, and the good one. Scored against thirteen labeled incidents from the same demo — eight failures, five of them run twice — ranger gets top-1 38% and names the wrong service in 23% of cases under this ranking mode. It is silent for the rest, and the silence is the half worth having.

Every case that ran twice agreed with itself, including both that fail, so those failures are systematic rather than unlucky. One case is a control where load rises and nothing breaks: declining is the correct answer there, and without it the decline column could not be read at all. The full accounting, including the cases excluded for having no defensible label, is in the README — and the reasoning behind the measurement is in the write-up.

github.com/Grace/ranger · Go · pre-alpha, and the README says so