I’m interested in designing ethical and accountable systems whose behavior can be explained because, "Trust, but verify," only works when trust is warranted. AI agents and LLMs are inherently untrustworthy: it's important for AI reasoning and actions to be understood, governed, and traced in real-time and after the fact.
I build production infrastructure, observability systems, and tools for working with them. I care about capturing the space between a system doing something and being able to establish what it did. I build tools for distributed tracing, telemetry normalization across incompatible instrumentation, AI infrastructure for model routing, replaying history, recording usage, reconciling it against billing, and enforcing explicit resource limits.
Much of my engineering work comes down to making the invisible visible, reducing mean time to recovery when systems break, and teaching systems how to gracefully self-heal.
Writing
-
Three places to put a translator
17 September 2026
"We are not maintaining a fork of the Collector so one vocabulary of spans can be renamed" is a correct objection, and it ends most conversations about a processor. The way out is to stop shipping a binary and start shipping the decision.
-
There is no model in here
11 September 2026
How normalizing GenAI telemetry actually works, in pseudocode: counting evidence instead of predicting it, an integer margin instead of a confidence score, and a rule table small enough to print. The highest score usually loses, and that is the interesting part.
-
Your telemetry was rewritten on the way here
8 September 2026
Collectors normalize it, backends remap it at ingest, schema files convert it between versions. All of that is load-bearing and none of it leaves a trace. There is no standard way for telemetry to say it was translated, or what the translation cost.
-
Don’t invent a language for this
8 September 2026
The obvious next step after normalizing LLM telemetry is a config format for the mappings, and a standard to submit it to. I tried the format. It was the wrong abstraction. The standard is worth doing — for something other than the mappings.
-
LLM telemetry has no standard. Here’s how to normalize it.
7 September 2026
Every GenAI instrumentation library spells the same facts differently, so nothing about your model calls can be queried as one thing — and the semantic conventions that would settle it currently have no released version to normalize against.
-
Duration is why every service on the path looks guilty
5 September 2026
A span's duration includes the time its children were running. When something deep in a call tree slows down, every ancestor slows by the same amount — and a ranked list of what got slower names all of them.
-
Your observability tool cannot be your evidence
5 September 2026 · Honeycomb said it themselves in 2023
Sampling is the product and retention is 60 days. What
finishing the other tier actually takes, and the sampler-ordering trap that
quietly makes an archive worthless.
-
You cannot filter your way to compliance
5 September 2026 · EU AI Act
The control most vendors sell is the one least able to
deliver what the regulation asks. Measured: past a single word substitution,
fuzzy matching falls off a cliff — and résumé lines score like attacks.
-
The strongest test in your AI controls checklist doesn't run
5 September 2026 · SOC 2
Reconciling logged requests against provider invoices is
the best test in every AI control mapping, including mine. AWS billing data
contains no request counts — here is what to reconcile instead.
All writing →
Demos
Most of these are the real Go tool compiled to WebAssembly and
running in your browser — not a reimplementation. The ranger demo is a report
the Go tool generated, with the incident data embedded. Nothing is uploaded.
-
ranger — find the service that actually broke
root-cause localization · runs in your browser
A real incident from the OpenTelemetry Demo, localized to
one service out of 87 ranked operations. Every ancestor on the request path
slowed by the same amount; self time is what separates the cause from the
callers waiting on it.
-
paladin — see what your agent reads
integrity monitoring · runs in your browser
Paste a CLAUDE.md or .cursorrules
and watch the characters your font refuses to draw appear: zero-width spaces,
bidi overrides that reorder text on screen, Cyrillic letters wearing Latin
faces.
-
warden — what a decline is allowed to say
rules engines · runs in your browser
Switch between a public, partner and internal caller and
watch what crosses the boundary change, with everything held back listed
underneath.
-
genai-interlingua — normalize a GenAI span
OpenTelemetry · runs in your browser
Things I've built
-
genai-interlingua
Go · OpenTelemetry processor · try it in your browser →
Six GenAI dialects, six names for the same token count,
so nothing about your model calls can be queried as one thing. Normalizes
them into one gen_ai.* schema and records what each translation
cost — four of seven captures found a mapping bug a completely green test
suite could not see.
-
ranger
Go · distributed tracing · try it in your browser →
Walks the OpenTelemetry trace DAG for the deepest span
whose deviation from baseline its children do not explain — the measurement
that separates the service that broke from every caller waiting on it.
Deterministic, and pre-alpha honestly: top-1 of 38% against thirteen labeled
incidents — eight distinct failures, five of them run twice — ranking roughly
ninety operations per window, with the confound that makes that a floor
published alongside it.
-
paladin
Go · integrity monitoring · try it in your browser →
Watches the files an AI coding agent obeys — CLAUDE.md,
.cursorrules, settings.json, skills — and reports drift, including the
zero-width characters, bidi overrides and homoglyphs a reviewer's font will
not draw.
-
warden
Go · rules engines · try it in your browser →
A guard between a rules engine and everyone else. A
declined transaction answers in the engine's own vocabulary — rule names,
working-memory facts, the threshold that priced the decision — and handing
that outward couples callers to names you can no longer change, leaks facts
that were never meant to leave, and turns a decline into an oracle a
fraudster can query. warden decides what crosses the boundary, and in what
vocabulary.
-
@gracefulcode/paladin
Node · supply chain
Audits an npm package before you install it: install
hooks, shipped agent instruction files, hidden characters, and whether it
carries a provenance attestation. Published through trusted publishing, with
no token in the repo.
Before this
I built software for research groups at MIT Lincoln Laboratory,
MIT CSAIL, UC Berkeley and UNC
Charlotte. At Lincoln Laboratory that meant
data visualization and logistics web interfaces — prototypes for
U.S. Department of Defense and USAMRDC programs, in Java, Ruby on
Rails and JavaScript — and presented them directly to U.S. Army stakeholders,
work that contributed to securing a million-dollar contract.
Elsewhere: multi-camera capture and depth reconstruction via stereoscopic animation at CSAIL,
RNA-sequencing visualization at Berkeley, and humanoid robot teleoperation driven
by optical hand tracking at UNC Charlotte.
A computer science degree with a focus in data visualization and computer
graphics, and an outside concentration in bioinformatics — which turns out to be
a reasonable education for thinking about pipelines.
GitHub · grace@gracefulco.de