<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Grace L. Christenbery's Blog</title>
  <link href="https://grace.github.io/feed.xml" rel="self"/>
  <link href="https://grace.github.io/"/>
  <updated>2026-09-17T19:30:16-04:00</updated>
  <id>https://grace.github.io/</id>
  <author><name>Grace L. Christenbery's Blog</name></author>
  
  <entry>
    <title>Three places to put a translator</title>
    <link href="https://grace.github.io/2026/09/17/three-places-to-put-a-translator/"/>
    <updated>2026-09-17T16:20:00-04:00</updated>
    <id>https://grace.github.io/2026/09/17/three-places-to-put-a-translator/</id>
    <content type="html">&lt;p&gt;Every previous post here argued about what a translated span should say. This
one is about a duller question that decides whether any of it gets used: where
does the translator actually run?&lt;/p&gt;

&lt;p&gt;The objection I keep hearing, and which is correct, goes like this. A platform
team has settled on &lt;code&gt;otelcol-contrib&lt;/code&gt;. It is in their base image, it is pinned,
it is in their upgrade rota, and somebody has already argued about which
processors are enabled. Asking them to maintain a custom Collector build so one
family of spans can be renamed is asking for a permanent maintenance obligation
in exchange for a dashboard. The answer is no, and it should be no.&lt;/p&gt;

&lt;p&gt;So there are three places this can run, and &lt;a href=&quot;/demos/genai-interlingua/architecture.html&quot;&gt;the architecture page&lt;/a&gt; draws
all three. Two are ordinary. The third is the one worth talking about.&lt;/p&gt;

&lt;h2 id=&quot;a-in-the-path&quot;&gt;A: in the path&lt;/h2&gt;

&lt;p&gt;Build &lt;code&gt;processor/genaiinterlingua&lt;/code&gt; into your Collector distribution and put it
between the receiver and the exporter. This is the version with the maintenance
obligation, and it is the only one that gets to normalize spans the moment they
arrive. It is a separate Go module, so building it in does not drag the CLI’s
dependencies along, and the core has none.&lt;/p&gt;

&lt;h2 id=&quot;b-beside-the-path&quot;&gt;B: beside the path&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;interlingua&lt;/code&gt; over captured OTLP JSON. This is how fixtures get built, how the
&lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/conformance.md&quot;&gt;conformance census&lt;/a&gt; is regenerated, and how you find out what a library
actually emits rather than what its documentation claims. Five of seven captures
found something the docs-derived version had wrong, which is the entire reason
the fixtures are captured and not written.&lt;/p&gt;

&lt;h2 id=&quot;c-not-in-the-path-at-all&quot;&gt;C: not in the path at all&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;-emit ottl&lt;/code&gt; prints a config for &lt;code&gt;transformprocessor&lt;/code&gt;. &lt;code&gt;-emit schema-file&lt;/code&gt;
prints a Telemetry Schema File for &lt;code&gt;schemaprocessor&lt;/code&gt;. Both of those processors
are already in the binary that platform team is running.&lt;/p&gt;

&lt;p&gt;So the mapping runs, in their pipeline, on their spans, under their upgrade
policy, and my code is not there. It ran once, on a laptop, to produce a text
file they read before applying.&lt;/p&gt;

&lt;p&gt;That is not a trick. It is what happens when the thing you have to distribute is
a decision rather than an implementation.&lt;/p&gt;

&lt;h2 id=&quot;why-it-works&quot;&gt;Why it works&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;normalize.Span&lt;/code&gt; does not mutate a span. It returns a &lt;code&gt;Result&lt;/code&gt;: attributes to
set, attributes to remove, which dialect claimed the span, how big the winner’s
margin was, and what the translation could not carry. A description of an edit,
not the edit.&lt;/p&gt;

&lt;p&gt;The Collector processor takes that description and applies it to pdata. The CLI
applies the same description to OTLP JSON. The emitters render it as OTTL
statements or as schema-file transformations for somebody else to apply. Three
appliers, one decision, and a disagreement between them would be a bug rather
than a configuration difference.&lt;/p&gt;

&lt;p&gt;The same property is why the mapping can be golden tested at all. You cannot
diff a mutation. You can diff a &lt;code&gt;Result&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;what-c-gives-up&quot;&gt;What C gives up&lt;/h2&gt;

&lt;p&gt;I would rather measure this than wave at it. A Telemetry Schema File carries
less than the processor does, because the schema-file format can express a
rename and cannot express reassembling a nested message document out of indexed
keys. The size of that shortfall is &lt;a href=&quot;/demos/genai-interlingua/conformance.html&quot;&gt;measured on the conformance page&lt;/a&gt;,
per dialect, rather than described as “some limitations apply.”&lt;/p&gt;

&lt;p&gt;The honest summary is that row C is right for the renames, which is most of the
volume, and row A is right when you need the parts of the mapping that are more
than a rename. Knowing which one you are in is a question about your own
telemetry, and the CLI in row B is how you answer it without deploying anything.&lt;/p&gt;

&lt;h2 id=&quot;the-part-i-actually-care-about&quot;&gt;The part I actually care about&lt;/h2&gt;

&lt;p&gt;There is a general rule hiding in here, and it took me a while to see it.&lt;/p&gt;

&lt;p&gt;A component that can only run as itself has to win a deployment argument before
it can be evaluated. A component that can emit its decision as somebody else’s
configuration gets evaluated first and deployed only if it earns it. The second
kind is much harder to say no to, and it is also much harder to build, because
it forces the decision to be data all the way through rather than a pile of
&lt;code&gt;if&lt;/code&gt; statements with side effects.&lt;/p&gt;

&lt;p&gt;It is also why the differences between this and the &lt;a href=&quot;https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/processor/genainormalizerprocessor&quot;&gt;upstream
&lt;code&gt;genainormalizer&lt;/code&gt;&lt;/a&gt; are going &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/upstream.md&quot;&gt;upstream&lt;/a&gt; rather than into a
competing component. If the useful part is the decision and not the binary, then
the decision belongs where the most people can run it, which is not here.&lt;/p&gt;

</content>
  </entry>
  
  <entry>
    <title>There is no model in here</title>
    <link href="https://grace.github.io/2026/09/11/there-is-no-model-in-here/"/>
    <updated>2026-09-11T20:30:00-04:00</updated>
    <id>https://grace.github.io/2026/09/11/there-is-no-model-in-here/</id>
    <content type="html">&lt;p&gt;Three posts ago I argued that &lt;a href=&quot;/2026/09/07/normalizing-llm-telemetry/&quot;&gt;GenAI instrumentation libraries each spell the
same facts differently&lt;/a&gt;, that &lt;a href=&quot;/2026/09/08/dont-invent-a-language-for-this/&quot;&gt;the mappings should be data but not a new
language&lt;/a&gt;, and that &lt;a href=&quot;/2026/09/08/what-did-the-translation-cost/&quot;&gt;a translated span should say what the translation
cost&lt;/a&gt;. Those were arguments. This one is the machine: what actually
happens to a span, stage by stage, in pseudocode.&lt;/p&gt;

&lt;p&gt;You can follow it without reading any Go, and there is no machine learning in it
anywhere. That is worth saying at the top, because “figure out which library
wrote this” sounds like a classification problem and gets built like one
depressingly often. It is a counting problem. Everything below follows from
that.&lt;/p&gt;

&lt;p&gt;Two pages run the real thing in your browser if you would rather poke at it than
read: the &lt;a href=&quot;/demos/genai-interlingua/&quot;&gt;normalizer&lt;/a&gt;, the &lt;a href=&quot;/demos/genai-interlingua/inspector/&quot;&gt;inspector&lt;/a&gt; for one span attribute at a
time, and &lt;a href=&quot;/demos/genai-interlingua/conformance.html&quot;&gt;what translation costs&lt;/a&gt; for the measurements across six
libraries.&lt;/p&gt;

&lt;h2 id=&quot;the-problem-in-four-lines&quot;&gt;The problem, in four lines&lt;/h2&gt;

&lt;p&gt;A model call produced 412 input tokens. Here are four of the ways it gets
written down:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;gen_ai.usage.prompt_tokens   = 412     OpenLLMetry
llm.token_count.prompt       = 412     OpenInference
ai.usage.promptTokens        = 412     Vercel AI SDK
gen_ai.usage.input_tokens    = 412     the conventions themselves
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;One fact, four spellings, and none of them can be summed with the others. So the
job is to rewrite all of them into one vocabulary. To rewrite a span you first
have to know which library wrote it, because nobody stamps their own name on it.&lt;/p&gt;

&lt;h2 id=&quot;stage-1-which-library-wrote-this&quot;&gt;Stage 1: which library wrote this?&lt;/h2&gt;

&lt;p&gt;Each library gets to look at the span and say how much evidence it sees &lt;em&gt;that
only it would have written&lt;/em&gt;. Not “does this look like an LLM call” — everyone
looks like an LLM call. Only the fingerprints.&lt;/p&gt;

&lt;p&gt;Braintrust’s entire answer is two lines, and it is the clearest one:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;score(span):
    if any attribute starts with &quot;braintrust.&quot;   →  3
    otherwise                                   →  0
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Most are a little richer. A whole namespace is worth 2 or 3; an individual
telltale key is worth 1:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;score_openllmetry(span):
    points = 0
    if span has any &quot;traceloop.*&quot;             points += 2
    for key in [gen_ai.usage.prompt_tokens,
                gen_ai.usage.completion_tokens,
                gen_ai.usage.total_tokens,
                llm.request.type]:
        if span has key                       points += 1
    if span has &quot;gen_ai.prompt.*&quot;
            or  &quot;gen_ai.completion.*&quot;         points += 1
    return points
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Then you pick the winner:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;detect(span):
    high, second, winner = 0, 0, none
    for each dialect:
        n = dialect.score(span)
        if n &amp;gt; high:     winner, high, second = dialect, n, high
        elif n &amp;gt; second: second = n

    confidence = high - second
    return winner, confidence
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;&lt;code&gt;high - second&lt;/code&gt; is the whole confidence model. If one library scored 7 and the
next best scored 0, the margin is 7. If they scored 5 and 3, the margin is 2 —
same winner, much less comfortable. That integer gets written onto the span as
&lt;code&gt;interlingua.dialect.confidence&lt;/code&gt;, and a human or a query can use it to find the
spans where the answer was close.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not a probability, and it cannot be turned into one.&lt;/strong&gt; There is no
denominator. The scores are unbounded evidence counts on scales that differ per
library — one library’s 3 is a namespace, another’s 3 is three separate keys —
so there is nothing to divide by. A percentage here would be a number invented
to look like a measurement, which is why &lt;a href=&quot;/demos/genai-interlingua/conformance.html&quot;&gt;the page that draws these
measurements&lt;/a&gt; does not publish one.&lt;/p&gt;

&lt;h3 id=&quot;the-highest-score-usually-loses&quot;&gt;The highest score usually loses&lt;/h3&gt;

&lt;p&gt;Here is the part I did not expect when I built it. Five of the six dialects are
named libraries. The sixth, &lt;code&gt;raw&lt;/code&gt;, recognizes hand-rolled instrumentation by
&lt;em&gt;shape&lt;/em&gt; — it counts attributes the conventions themselves define, plus a table of
common folk spellings:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;score_raw(span):
    conformant = keys the conventions define        # worth 2 each
    folk       = keys in the folk-spelling table    # worth 1 each

    if conformant == 0 and folk &amp;lt; 2:  return 0      # a lone `model` attribute
    return conformant*2 + folk                      # is evidence of nothing
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Because it counts &lt;em&gt;every&lt;/em&gt; conventions-defined key at 2, &lt;code&gt;raw&lt;/code&gt; scores enormously
on spans that belong to somebody else. On a real LangChain span:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;span&lt;/th&gt;
      &lt;th&gt;raw&lt;/th&gt;
      &lt;th&gt;openllmetry&lt;/th&gt;
      &lt;th&gt;winner&lt;/th&gt;
      &lt;th&gt;margin&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;ChatOpenAI.chat&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;32&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;3&lt;/td&gt;
      &lt;td&gt;openllmetry&lt;/td&gt;
      &lt;td&gt;3&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;litellm_request&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;24&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;3 (litellm 5)&lt;/td&gt;
      &lt;td&gt;litellm&lt;/td&gt;
      &lt;td&gt;2&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;ai.generateText&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;0 (vercel 7)&lt;/td&gt;
      &lt;td&gt;vercel&lt;/td&gt;
      &lt;td&gt;7&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;chat gpt-4o-mini&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;26&lt;/td&gt;
      &lt;td&gt;0&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;raw&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;code&gt;raw&lt;/code&gt; out-scores the winner on eight of the fifteen spans in the corpus and wins
none of the ones it competes for. It is marked as a &lt;strong&gt;fallback&lt;/strong&gt; and excluded
from the race entirely — consulted only when nothing that knows what it is
looking at has claimed the span:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;detect(span):
    contenders = dialects that are not fallbacks
    ...argmax over contenders, margin = high - second...

    if nobody claimed it:
        winner = best-scoring fallback
        confidence = 0            # it did not beat anything.
                                  # it was the only thing left.
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The reason this belongs in detection rather than in the scoring is precise, and
it is the nicest small design argument in the codebase: a fallback that scores on
every GenAI span &lt;em&gt;would not win&lt;/em&gt; those spans, but the points it took would come
straight out of the winner’s margin. A positive identification would read as less
confident merely because a fallback also exists. So the fallback tier is
separate, and a span labeled &lt;code&gt;raw&lt;/code&gt; with a confidence of 0 is telling the truth
about how it was identified.&lt;/p&gt;

&lt;p&gt;Note the last row. Margin 0 does not mean “coin flip”. It means “claimed from the
fallback tier”.&lt;/p&gt;

&lt;h2 id=&quot;stage-2-translating-it&quot;&gt;Stage 2: translating it&lt;/h2&gt;

&lt;p&gt;Now the winner reads the span. About half of that work is a lookup table. A rule
is three things:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;Rule:
    field      the concept being produced        (not an attribute name)
    keys       source keys, in precedence order  (first one present wins)
    transform  what happens to the value on the way through
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The &lt;code&gt;field&lt;/code&gt; is deliberately not a name. &lt;code&gt;gen_ai.usage.cache_write.input_tokens&lt;/code&gt;
in the current draft of the conventions is spelled
&lt;code&gt;gen_ai.usage.cache_creation.input_tokens&lt;/code&gt; in the released version, and it is the
same idea both times. So the middle of the pipeline traffics in concepts, and
names are applied last, on the way out.&lt;/p&gt;

&lt;p&gt;The transform is a &lt;strong&gt;closed set&lt;/strong&gt; — six operations, applied in a fixed order:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;transform:
    lower         lowercase it            &quot;OpenAI&quot; → &quot;openai&quot;
    cut_after     keep the part before a separator
    trim_suffix   strip a known suffix
    map           look it up in a table   (+ what to do if it&apos;s not in the table)
    divide        unit conversion         milliseconds → seconds
    to_list       wrap a scalar in a one-element list
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Here is a real rule, and it uses three of them. The Vercel AI SDK writes its
provider as &lt;code&gt;&amp;lt;provider&amp;gt;.&amp;lt;api surface&amp;gt;&lt;/code&gt; — &lt;code&gt;openai.chat&lt;/code&gt;,
&lt;code&gt;amazon-bedrock.messages&lt;/code&gt;, &lt;code&gt;google.generative-ai&lt;/code&gt; — and only the first segment
names the provider:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;{ field: provider_name,
  keys:  [ai.model.provider, gen_ai.system],
  transform: lower, cut_after &quot;.&quot;, map(vercel_providers) }
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;&lt;strong&gt;Every one of those six operations is a total function of one value.&lt;/strong&gt; It cannot
look at another attribute, iterate the span, or branch on anything but its own
input. That constraint is the whole reason the table is a table: a rule that can
be &lt;em&gt;stated&lt;/em&gt; can also be printed, audited, and exported — into an OpenTelemetry
Collector config that runs without this binary at all.&lt;/p&gt;

&lt;p&gt;And it is why the other half of the mapping is not a table. Reassembling
&lt;code&gt;gen_ai.prompt.0.tool_calls.1.name&lt;/code&gt;-style indexed attributes into one nested
document, or deciding what &lt;code&gt;traceloop.entity.name&lt;/code&gt; means by looking at a sibling
attribute, are not transforms of a value. They are &lt;em&gt;readings of a span&lt;/em&gt;. Those
stay as code, and the tool reports them as the part an export cannot carry. I
&lt;a href=&quot;/2026/09/08/dont-invent-a-language-for-this/&quot;&gt;tried making them data&lt;/a&gt; and that is the post about why it failed.&lt;/p&gt;

&lt;h2 id=&quot;stage-3-what-could-not-be-carried&quot;&gt;Stage 3: what could not be carried&lt;/h2&gt;

&lt;p&gt;This is the part other normalizers do not do, and it is the only reason this
project is interesting.&lt;/p&gt;

&lt;p&gt;When a rule cannot carry something, the reason is recorded on the span rather
than dropped. Five reasons, and they are five genuinely different problems:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;reason&lt;/th&gt;
      &lt;th&gt;what it means&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;no_field&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;the conventions have &lt;strong&gt;no concept&lt;/strong&gt; for this, at any version&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;unstructured&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;a provider-shaped blob the translator declined to parse&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;flattened&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;indexed attributes collapsed into one field; per-index structure lost&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;coerced&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;the value survived but its type did not&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;ambiguous&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;it could have meant two different fields, so nothing was guessed&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;code&gt;coerced&lt;/code&gt; has the best small example. The &lt;code&gt;divide&lt;/code&gt; operation exists because the
Vercel SDK reports milliseconds where the conventions specify seconds. If the
value is not a number, it is &lt;strong&gt;not&lt;/strong&gt; passed through unchanged:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;divide(value, by):
    if value is not a number:
        return loss(coerced, &quot;value is not a number&quot;)
    return value / by
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Passing it through would put a millisecond count under an attribute the
conventions define in seconds — wrong by exactly a thousand, and it looks fine.
Refusing is the more useful answer.&lt;/p&gt;

&lt;p&gt;(A pedantic detail I enjoy: it divides by 1000 rather than multiplying by 0.001.
Over the first 200,000 millisecond values those two disagree in the last bit
26,651 times, because 0.001 has no exact float64 representation and 1000 does.)&lt;/p&gt;

&lt;h3 id=&quot;the-counter-that-counts-what-you-forgot&quot;&gt;The counter that counts what you forgot&lt;/h3&gt;

&lt;p&gt;Recording losses as you find them has an obvious failure: you only record the
ones you thought to name. So after the rules run, there is a sweep. For every
attribute under a namespace this dialect &lt;em&gt;claimed as its own&lt;/em&gt;, that it neither
read nor already reported, the sweep names it and records &lt;code&gt;no_field&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Without that, &lt;code&gt;lossy.count = 0&lt;/code&gt; means “I recognized none of the things I know I
drop,” which is very nearly the opposite of what any reader would take it for. A
loss counter that only counts remembered losses is not a measurement. The
shape-matching fallback had this right first and the five purpose-built dialects
did not, which is an embarrassing way round for it to be.&lt;/p&gt;

&lt;h3 id=&quot;what-that-measures-across-six-real-dialects&quot;&gt;What that measures, across six real dialects&lt;/h3&gt;

&lt;p&gt;Six dialects, nine span fixtures — seven captured from the libraries running,
two hand-built — and thirty distinct fields. &lt;strong&gt;105 source attributes have
nowhere in the conventions to go:&lt;/strong&gt;&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;dialect&lt;/th&gt;
      &lt;th&gt;fields carried (of 30)&lt;/th&gt;
      &lt;th&gt;attributes with no home&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;openllmetry&lt;/td&gt;
      &lt;td&gt;21&lt;/td&gt;
      &lt;td&gt;24&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;vercel&lt;/td&gt;
      &lt;td&gt;19&lt;/td&gt;
      &lt;td&gt;8&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;openinference&lt;/td&gt;
      &lt;td&gt;17&lt;/td&gt;
      &lt;td&gt;7&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;litellm&lt;/td&gt;
      &lt;td&gt;13&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;53&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;raw&lt;/td&gt;
      &lt;td&gt;13&lt;/td&gt;
      &lt;td&gt;6&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;braintrust&lt;/td&gt;
      &lt;td&gt;9&lt;/td&gt;
      &lt;td&gt;7&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Ninety of the 105 are &lt;code&gt;no_field&lt;/code&gt; — not a missing name, a missing &lt;em&gt;concept&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;LiteLLM’s 53 is not noise and it is not sloppiness. It is an entire cost model —
&lt;code&gt;gen_ai.cost.input_cost&lt;/code&gt;, &lt;code&gt;output_cost&lt;/code&gt;, &lt;code&gt;margin_percent&lt;/code&gt;, &lt;code&gt;discount_amount&lt;/code&gt; —
plus every &lt;code&gt;user_api_key_*&lt;/code&gt; field recording who spent what. The conventions price
nothing, at any version. &lt;strong&gt;That is a finding about the conventions, not about
LiteLLM&lt;/strong&gt;, and costing is a large part of why anyone puts a proxy in front of a
model in the first place.&lt;/p&gt;

&lt;p&gt;One more thing about the 105, because it is the question everyone asks: nothing
was deleted. Under the default, every one of those attributes is still on the
span under the name its library gave it. What it lacks is a &lt;code&gt;gen_ai.*&lt;/code&gt; name, so a
query written in the conventions’ vocabulary will not find it. The tool has a
mode that removes them and it is not the default, because that is a choice
somebody should make on purpose.&lt;/p&gt;

&lt;h2 id=&quot;stage-4-pinning-a-target-costs-something-too&quot;&gt;Stage 4: pinning a target costs something too&lt;/h2&gt;

&lt;p&gt;There is a second translation, and a second kind of loss. Concepts have to be
written out under some specific version of the conventions, and the versions
differ:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;render(field, value, target):
    key = target.name_for(field)
    if no key:            return loss(no_attribute)   # this release has no such attribute
    if not target.accepts(field, value):
        return loss(no_value)                          # it has the attribute, not this value
    write key = value
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The released &lt;code&gt;v1.41.0&lt;/code&gt; can express 27 of the 30 fields; the current draft can
express all 30. Concretely, pinning the released version costs you
&lt;code&gt;gen_ai.prompt.version&lt;/code&gt; and both audio-token fields. And &lt;code&gt;no_value&lt;/code&gt; bites on both
targets: OpenLLMetry states an operation name and a provider name that are
outside the conventions’ permitted value sets, so they are recorded as refused
rather than written.&lt;/p&gt;

&lt;p&gt;Both choices are defensible. Neither is correct. What is not defensible is making
the choice silently, which is what a normalizer with a hardcoded table does —
&lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/moving-target.md&quot;&gt;so the target is a flag&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;what-ends-up-on-the-span&quot;&gt;What ends up on the span&lt;/h2&gt;

&lt;p&gt;The translation writes its own account of itself, in attributes you can group by:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;interlingua.dialect            = openllmetry
interlingua.dialect.confidence = 7
interlingua.target             = v1.41.0
interlingua.lossy              = [gen_ai.completion.*, gen_ai.prompt.version, ...]
interlingua.lossy.count        = 4
interlingua.hops               = 1
interlingua.mapping            = 9f2c1ab4e07d5c83
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Three choices there that took a while to get right.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;lossy.count&lt;/code&gt; is written &lt;strong&gt;even when it is zero&lt;/strong&gt;, because “this span lost
nothing” and “this span was never normalized” are different facts and several
backends store an array attribute as an opaque string you cannot count.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;hops&lt;/code&gt; is an integer rather than a chain of translations. That is a concession to
reality: Honeycomb types every array-valued attribute as a string column —
including the conventions’ own &lt;code&gt;gen_ai.input.messages&lt;/code&gt; — so a chain would arrive
as a blob you cannot group by, filter into, or count. An integer you can alert on
beats a structure you cannot query.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;mapping&lt;/code&gt; is the one I would defend hardest. It is a digest over &lt;strong&gt;the rule
tables themselves&lt;/strong&gt; — every rule, every transform field including the ones left
at their defaults, every target’s key map — and not a commit hash. A commit hash
moves on every commit, including the ones that change nothing a span could
possibly see, so it cannot answer “was this span translated under the same rules
as that one?” The criterion is worth stating as a general rule, because it is the
same one that makes me refuse a conformance percentage:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;A number is a measurement only if it moves when the measured thing moves, and
not otherwise.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The digest’s honest limit, stated rather than papered over: rewriting &lt;em&gt;how&lt;/em&gt; a
method reassembles messages, while changing nothing about which fields come out,
leaves the digest where it was.&lt;/p&gt;

&lt;h2 id=&quot;go-and-look&quot;&gt;Go and look&lt;/h2&gt;

&lt;p&gt;Everything above is running in your browser, compiled from the same Go the
command line and the Collector processor run, so the pages cannot disagree with
the tool about what a span becomes. Nothing you paste is uploaded; there is no
server to upload it to.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;&lt;a href=&quot;/demos/genai-interlingua/&quot;&gt;Normalize a span&lt;/a&gt;&lt;/strong&gt; — paste one in, or pick a real capture, and watch
it rewritten.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;a href=&quot;/demos/genai-interlingua/inspector/&quot;&gt;The inspector&lt;/a&gt;&lt;/strong&gt; — one attribute at a time: which source key produced
this value, what evidence the reading turned on, and what it could not carry.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;a href=&quot;/demos/genai-interlingua/conformance.html&quot;&gt;What translation costs&lt;/a&gt;&lt;/strong&gt; — the measurements above, drawn: the scoring
race per span, and where all 105 losses go.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;a href=&quot;https://github.com/Grace/genai-interlingua&quot;&gt;source is here&lt;/a&gt;. Two of its documents are generated by the test suite
rather than written, which is the only reason I trust the numbers in this post:
&lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/conformance.md&quot;&gt;the census&lt;/a&gt; is regenerated from the fixtures and CI fails when it and
the fixtures disagree, and &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/findings.md&quot;&gt;what capturing the real libraries found&lt;/a&gt; is
the log of the five defects that turned up when I stopped writing fixtures from
documentation and started recording what the libraries actually emit. Five of
seven captures found something wrong. Every one of them was invisible to a
completely green test suite.&lt;/p&gt;

&lt;h2 id=&quot;what-this-is-not&quot;&gt;What this is not&lt;/h2&gt;

&lt;p&gt;There is already a first-party component for this. &lt;a href=&quot;https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/processor/genainormalizerprocessor&quot;&gt;&lt;code&gt;processor/genainormalizer&lt;/code&gt;&lt;/a&gt;
has been in &lt;code&gt;opentelemetry-collector-contrib&lt;/code&gt; since February and ships in the
&lt;code&gt;otelcol-contrib&lt;/code&gt; binary you are probably already running. It maps OpenInference
and OpenLLMetry and needs no custom build. &lt;strong&gt;If those are the libraries you have,
reach for it first.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the claim here is not that translating GenAI telemetry is novel. It is that
none of the translators record what the translation cost. That component’s own
source says &lt;code&gt;callers must drop the attribute&lt;/code&gt;, and with &lt;code&gt;remove_originals: true&lt;/code&gt;
the source attribute goes too — so the evidence leaves with the data. Those 105
attributes are what that silence is hiding, and the differences are going
&lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/upstream.md&quot;&gt;upstream&lt;/a&gt; rather than into a competing component, because a second
one would be worse for everyone than one good one.&lt;/p&gt;

&lt;p&gt;Normalization you cannot audit is just another assertion. Everything in this post
exists to make the assertion checkable.&lt;/p&gt;

</content>
  </entry>
  
  <entry>
    <title>Your telemetry was rewritten on the way here</title>
    <link href="https://grace.github.io/2026/09/08/what-did-the-translation-cost/"/>
    <updated>2026-09-08T02:08:31-04:00</updated>
    <id>https://grace.github.io/2026/09/08/what-did-the-translation-cost/</id>
    <content type="html">&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Added 2026-09-09.&lt;/strong&gt; When I wrote this I was arguing from my own normalizer.
It turns out the strongest evidence for the argument is in someone else’s:
&lt;a href=&quot;https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/processor/genainormalizerprocessor&quot;&gt;&lt;code&gt;processor/genainormalizer&lt;/code&gt;&lt;/a&gt;, which ships in &lt;code&gt;otelcol-contrib&lt;/code&gt;, drops
what it cannot carry — by design, and it says so in its own source. From
&lt;code&gt;internal/otelsemconv/coerce.go&lt;/code&gt;: &lt;em&gt;“src cannot be safely coerced; callers must
drop the attribute”&lt;/em&gt; and &lt;em&gt;“Map / Slice / Bytes: do not stringify. Caller drops
the rename.”&lt;/em&gt; With its &lt;code&gt;remove_originals: true&lt;/code&gt; the source attribute is deleted
too, so the evidence goes with the data. That is not a criticism of a component
I am now &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/upstream.md&quot;&gt;contributing to&lt;/a&gt;. It is the point of this post, written by
someone else, in a comment, where no query can reach it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here is a span attribute: &lt;code&gt;gen_ai.usage.input_tokens: 412&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;You cannot tell, from that, whether the library emitted it that way. It might
have been renamed from &lt;code&gt;gen_ai.usage.prompt_tokens&lt;/code&gt; by a collector. It might have
been translated from &lt;code&gt;llm.token_count.prompt&lt;/code&gt; by something that also dropped
three attributes it had nowhere to put. It might have been rewritten at ingest by
your vendor, into or out of a vocabulary you never chose. All four produce a span
that looks exactly like this one.&lt;/p&gt;

&lt;p&gt;That is not a hypothetical about a thing that might start happening. It is a
description of a Tuesday.&lt;/p&gt;

&lt;h2 id=&quot;everything-rewrites-telemetry&quot;&gt;Everything rewrites telemetry&lt;/h2&gt;

&lt;p&gt;Four layers, all doing it now, none of them exotic:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A collector normalizes vendor attributes.&lt;/strong&gt; This is the recommended practice.
It is what the &lt;code&gt;transform&lt;/code&gt; processor is for and what I spent &lt;a href=&quot;/2026/09/08/dont-invent-a-language-for-this/&quot;&gt;the last two
posts&lt;/a&gt; building.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A backend remaps at ingest.&lt;/strong&gt; Arize AX converts inbound &lt;code&gt;gen_ai.*&lt;/code&gt; attributes
into OpenInference before your spans are stored. That is documented and
defensible, and it means the attribute names you query are not the attribute
names you sent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A schema file converts between versions.&lt;/strong&gt; OpenTelemetry has a whole mechanism
for this: &lt;a href=&quot;https://opentelemetry.io/docs/specs/otel/schemas/&quot;&gt;telemetry schemas&lt;/a&gt;, applied in the collector, renaming
attributes as conventions evolve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A migration is bridged.&lt;/strong&gt; When the HTTP conventions renamed half their
attributes, everybody ran something in the middle for a year.&lt;/p&gt;

&lt;p&gt;Every one of these is a good idea. Together they mean the telemetry arriving in
your backend has been through an unknown number of rewrites, by an unknown set of
implementations, losing an unknown amount on the way — and there is no standard
way for any of them to say so.&lt;/p&gt;

&lt;h2 id=&quot;schema_url-answers-a-different-question&quot;&gt;&lt;code&gt;schema_url&lt;/code&gt; answers a different question&lt;/h2&gt;

&lt;p&gt;The nearest existing mechanism is &lt;code&gt;schema_url&lt;/code&gt;, and it is worth being precise
about why it does not cover this.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;schema_url&lt;/code&gt; says which version of a schema the telemetry &lt;em&gt;claims to conform
to&lt;/em&gt;. That is useful. It is not the same as saying a conversion happened, which
direction it went, what the source vocabulary was when the source was not a
schema version at all, or what the conversion could not carry.&lt;/p&gt;

&lt;p&gt;And it is absent exactly when it would help most. The GenAI semantic conventions
currently have no released version to point at — the attributes were deprecated
out of the main repository and moved to one that has never tagged anything — so
there is no URL to put there. That is not an edge case I went looking for; it is
the situation that made me build a normalizer in the first place.&lt;/p&gt;

&lt;h2 id=&quot;the-failure-this-produces&quot;&gt;The failure this produces&lt;/h2&gt;

&lt;blockquote&gt;
  &lt;p&gt;A translation layer that drops data silently is worse than no translation
layer, because you will trust the result.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence is the whole argument and I would like to defend it with something
other than assertion.&lt;/p&gt;

&lt;p&gt;While building the thing, I recorded on every rewritten span the list of keys
that span was &lt;em&gt;not&lt;/em&gt; a faithful carrier of. Not a log line — an attribute, next to
the data, queryable. It is the least glamorous feature in the repository.&lt;/p&gt;

&lt;p&gt;Five of seven captures of real instrumentation libraries turned up a mapping bug
that a completely green test suite could not see. Every one was found by reading
that list and asking why something was on it. The suite was green because the
tests encoded the same misunderstanding the code did. The loss list was the only
artifact that disagreed with me.&lt;/p&gt;

&lt;p&gt;Then it did it again, in the export path. A generated OTTL config that parsed
cleanly in a real collector turned out to carry less than its own header claimed,
and I only found out by running both implementations over the same spans and
diffing the results — which is, again, just the question “what did this
translation cost” asked mechanically.&lt;/p&gt;

&lt;p&gt;I am not claiming the idea is clever. I am claiming it is the only part of the
system that has ever told me I was wrong.&lt;/p&gt;

&lt;h2 id=&quot;what-to-record&quot;&gt;What to record&lt;/h2&gt;

&lt;p&gt;Six attributes. No file format changes, no SDK work, no new parser — a namespace
and a convention:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt; &lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;source&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;The vocabulary the telemetry arrived in. A schema URL when there is one; otherwise an identifier, because the common case is that the source declared nothing.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;source confidence&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;How decisive the identification was, when it was inferred rather than declared.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;target&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;What it was translated into.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;lossy&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;The keys this telemetry is &lt;em&gt;not&lt;/em&gt; a faithful carrier of.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;lossy count&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;The length of that list, &lt;strong&gt;written even when zero&lt;/strong&gt;.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;by&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;Which implementation did it, when several implementations of one mapping exist and they do not all carry the same amount.&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Two of those deserve defending, because both look redundant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The count, when it is zero.&lt;/strong&gt; “This translation lost nothing” and “nobody
checked” are different claims and an absent array cannot tell them apart. Writing
zero is an assertion. It is also the difference between a dimension your backend
will aggregate over and one it will not — “which service is losing the most in
translation” should be a query, not an afternoon.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The confidence.&lt;/strong&gt; In practice the source is almost never declared.
Instrumentation libraries do not stamp their own name on spans, so anything
handling more than one has to &lt;em&gt;infer&lt;/em&gt; which it is looking at from the shape of
the attributes. That inference is usually easy and occasionally a coin flip. An
inference recorded without its confidence is indistinguishable from a fact.&lt;/p&gt;

&lt;h2 id=&quot;what-this-deliberately-is-not&quot;&gt;What this deliberately is not&lt;/h2&gt;

&lt;p&gt;It is not a way to describe the mapping. That is a much larger proposal, it
&lt;a href=&quot;/2026/09/08/dont-invent-a-language-for-this/&quot;&gt;belongs to a different argument&lt;/a&gt;, and it is considerably less likely to
go anywhere.&lt;/p&gt;

&lt;p&gt;This one is only: telemetry that has been translated should be able to say so.
Worth doing even if no transformation format ever changes, because the
translations are happening right now, in collectors and OTTL configs and vendor
ingest pipelines, and not one of them can leave a trace.&lt;/p&gt;

&lt;h2 id=&quot;where-it-goes&quot;&gt;Where it goes&lt;/h2&gt;

&lt;p&gt;OpenTelemetry, and specifically the Semantic Conventions SIG — this is an
attribute convention, not a protocol change. Not W3C, who own trace &lt;em&gt;context&lt;/em&gt;,
the propagation format on the wire, and not attribute semantics. Not IETF.&lt;/p&gt;

&lt;p&gt;I have &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/oteps/0001-translation-provenance.md&quot;&gt;drafted it&lt;/a&gt; and &lt;strong&gt;not filed it&lt;/strong&gt;, which I would rather say plainly
than let a directory called &lt;code&gt;oteps/&lt;/code&gt; imply otherwise. The order that works for
this kind of thing is: ship it, get a couple of real users, open an issue
describing the problem rather than your solution, show up to the SIG call twice
before proposing anything, file small useful patches first. The document comes
last, if it comes at all. An OTEP filed cold by somebody with no history in a
project is a document that gets politely queued, and that is a reasonable thing
for a project to do with it.&lt;/p&gt;

&lt;p&gt;What exists today is the running version, under a namespace named after my own
tool — which is exactly what a shared convention must not be, and why the draft
proposes a neutral one. It is published as a &lt;a href=&quot;https://github.com/Grace/genai-interlingua/tree/main/registry&quot;&gt;Weaver registry&lt;/a&gt; that
&lt;code&gt;weaver registry check&lt;/code&gt; validates, so it can be resolved and depended on without
adopting a line of my code.&lt;/p&gt;

&lt;p&gt;That is a different conversation from a proposal that exists only as prose. It is
not yet the conversation, and I would rather be at the beginning of it honestly
than describe it as further along.&lt;/p&gt;

</content>
  </entry>
  
  <entry>
    <title>Don’t invent a language for this</title>
    <link href="https://grace.github.io/2026/09/08/dont-invent-a-language-for-this/"/>
    <updated>2026-09-08T00:51:03-04:00</updated>
    <id>https://grace.github.io/2026/09/08/dont-invent-a-language-for-this/</id>
    <content type="html">&lt;p&gt;&lt;a href=&quot;/2026/09/07/normalizing-llm-telemetry/&quot;&gt;Last post&lt;/a&gt; I argued that GenAI instrumentation libraries each spell the
same facts differently, that you should normalize them at the pipeline, and that
whatever your normalizer cannot carry should be written onto the span rather
than dropped in silence.&lt;/p&gt;

&lt;p&gt;Two questions came back, and they are the right two. Should the mappings become
a format — JSON, a small DSL, something SPL-shaped — instead of code? And should
any of this be standardized, by somebody other than me?&lt;/p&gt;

&lt;p&gt;No, and yes. The &lt;em&gt;no&lt;/em&gt; is more interesting than it sounds: I built the format
before deciding against it, and what killed it was not taste. And the &lt;em&gt;yes&lt;/em&gt;
turns out not to be about the mappings at all, which is where this ends up and
the part I would keep if you only read one section.&lt;/p&gt;

&lt;h2 id=&quot;three-layers-that-keep-getting-conflated&quot;&gt;Three layers that keep getting conflated&lt;/h2&gt;

&lt;p&gt;Normalizing telemetry is three separable things, and almost every conversation
about it slides between them mid-sentence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rules.&lt;/strong&gt; &lt;code&gt;llm.token_count.prompt&lt;/code&gt; and &lt;code&gt;gen_ai.usage.input_tokens&lt;/code&gt; are the
same fact. That is data. It is a table.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The execution.&lt;/strong&gt; Something has to actually rewrite the span, in a pipeline, at
volume, without dropping anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The claim.&lt;/strong&gt; An assertion that the two names &lt;em&gt;are&lt;/em&gt; the same fact, that
translating between them costs nothing — or costs exactly this much — and that
the assertion is auditable by somebody who was not there when it was made.&lt;/p&gt;

&lt;p&gt;The industry has quietly settled the second layer several times over, privately,
inside each vendor’s ingest. It has never published the third. Arize AX rewrites
inbound &lt;code&gt;gen_ai.*&lt;/code&gt; attributes into OpenInference before your spans are stored.
That is a defensible product decision and I am not picking on them; the point is
that you cannot read the mapping, cannot version it, and cannot find out what it
discarded. Every backend is doing some version of this and none of them will
show you the table.&lt;/p&gt;

&lt;p&gt;So: which layer would a new language be for?&lt;/p&gt;

&lt;h2 id=&quot;the-rules-should-be-data-that-is-not-the-same-as-a-new-format&quot;&gt;The rules should be data. That is not the same as a new format.&lt;/h2&gt;

&lt;p&gt;I went and did this to my own code, which is the only honest way to find out
what it costs.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;genai-interlingua&lt;/code&gt; had two halves that both looked like tables and only one of
which was. The schema half — which attribute name each convention version uses,
which values it admits — is now &lt;em&gt;generated&lt;/em&gt; from the registries OpenTelemetry
publishes, rather than transcribed by me. That was straightforwardly correct and
I should have done it first. The derivation turned out to be a rule with one
exception: the attribute key is &lt;code&gt;gen_ai.&lt;/code&gt; plus the field name, except for the
token-cache field that v1.41.0 spells &lt;code&gt;cache_creation&lt;/code&gt; and the new repository
spells &lt;code&gt;cache_write&lt;/code&gt;. Fifty keys and seventy-two keys, reproduced exactly. The
hand-typed tables had been right, which I would not have bet on.&lt;/p&gt;

&lt;p&gt;The mapping half is where it gets interesting. I converted five of the six
dialects from Go calls into declared tables — 86 mappings state cleanly as data
against 38 that do not, so roughly seven in ten. Stating one looks like: read
this key, lowercase it, cut it at the first dot, run it through this lookup,
write it there.&lt;/p&gt;

&lt;p&gt;The ratio is an average over emitters that are nothing like each other, which
matters more than the average. LiteLLM is 20 against 4. OpenInference is 16
against 16, because it packs every sampling parameter — temperature, top_p, the
penalties, the seed, nine others — into one JSON string, and no table reaches
inside a string. The sixth dialect is not migratable at all: it is the fallback
for spans nobody designed, and it works by resolving each attribute against the
entire semantic-convention registry at runtime, so writing it as data would
either duplicate the registry or admit it carries nothing.&lt;/p&gt;

&lt;p&gt;The three in ten that do not state are not waiting for a better format.&lt;/p&gt;

&lt;p&gt;Consider reassembling &lt;code&gt;gen_ai.prompt.0.tool_calls.1.arguments&lt;/code&gt; — two levels of
indexing — into one nested JSON document. Or deciding what &lt;code&gt;traceloop.entity.name&lt;/code&gt;
means by looking at whether a sibling attribute says this span is an agent or a
tool. Or refusing to map a Braintrust span carrying three evaluation scores,
because the conventions model one evaluation per span and picking one silently
would be worse than admitting the mismatch.&lt;/p&gt;

&lt;p&gt;Those are not transformations of a value. They are &lt;em&gt;readings of a span&lt;/em&gt; — they
need the whole span in hand, not one attribute, because what they produce
depends on what else is there. A format expressive enough to state them has
conditionals, loops and a JSON parser, at which point you have written a
programming language, and a worse one than the several already available.&lt;/p&gt;

&lt;h2 id=&quot;and-one-of-those-already-exists&quot;&gt;And one of those already exists&lt;/h2&gt;

&lt;p&gt;OTTL — the OpenTelemetry Transformation Language — is a DSL for exactly this,
governed by OpenTelemetry, shipped in every Collector distribution, running in
production at a scale mine never will. Proposing a new config format for
telemetry rewriting in 2026 means proposing to compete with it. That is a bad
trade for a mapping table.&lt;/p&gt;

&lt;p&gt;The better move is the opposite one: emit OTTL. So &lt;code&gt;genai-interlingua&lt;/code&gt; now does.&lt;/p&gt;

&lt;pre&gt;&lt;code class=&quot;language-console&quot;&gt;$ interlingua -emit ottl -dialect litellm -target v1.41.0
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Out comes a &lt;code&gt;transform&lt;/code&gt; processor you paste into your own Collector. No custom
build, no Go, no dependency on my repository continuing to exist.&lt;/p&gt;

&lt;p&gt;A stock Collector accepts it; I checked, in both target schemas, because a
generated config that has only ever been compared against a golden file has not
been tested, it has been photographed.&lt;/p&gt;

&lt;p&gt;Then I checked the thing that actually matters, which is whether it &lt;em&gt;means&lt;/em&gt; the
same thing. Every captured span goes through the Go processor and through a real
Collector running the exported config, and the two outputs are compared
attribute by attribute. It found four bugs on its first run, and every one of
them had been sitting behind a config that parsed cleanly and looked right:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;A scalar-to-list conversion that wrapped unconditionally, where the Go code
wraps only a string. On the one emitter whose value is already an array, the
attribute came out as a list containing a list.&lt;/li&gt;
  &lt;li&gt;A lookup table that never fired on array values at all. Go maps over the
elements; OTTL has no iteration, so the generated statements compare the whole
array against each table key and match nothing. That one is not fixable — it
is now declared instead, which is the point.&lt;/li&gt;
  &lt;li&gt;An enum guard that was &lt;em&gt;more&lt;/em&gt; destructive than the processor, deleting a
non-conformant value the processor deliberately preserves.&lt;/li&gt;
  &lt;li&gt;And a header claiming coverage it did not have.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The third and fourth are the ones I would have shipped. An export that quietly
does less than it advertises is the exact failure this post is about, and I
wrote it into my own tool while writing the argument against it.&lt;/p&gt;

&lt;p&gt;And because the export is a subset, it says so in its own header:&lt;/p&gt;

&lt;pre&gt;&lt;code class=&quot;language-yaml&quot;&gt;# WHAT THIS CARRIES (20 attributes)
# ...
# WHAT THIS CANNOT CARRY (4 attributes)
#
#   gen_ai.input.messages
#   gen_ai.output.messages
#   gen_ai.response.finish_reasons
#   gen_ai.tool.definitions
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Three things do not survive the trip. The unstated mappings, above. Per-span
loss accounting, because &lt;code&gt;interlingua.lossy&lt;/code&gt; is computed from what a given span
turned out to carry and a static config can only describe itself. And detection:
picking a dialect means scoring every candidate across the whole span and taking
the maximum, and OTTL has neither a loop nor an argmax. An exported config is
pinned to one library and scoped by the spellings only that library uses.&lt;/p&gt;

&lt;p&gt;That last one is a genuine limit, not a gap to fill later. It is also the
smallest of the three.&lt;/p&gt;

&lt;p&gt;The rule I would give anyone else: &lt;strong&gt;if your mapping can be stated as data,
state it as data, then hand the data to a language somebody else maintains.&lt;/strong&gt;
Keep code for the rules that are readings rather than transformations, and be
loud about which is which. My compiler cannot tell those apart, so a test does —
it runs every captured span through the parser and fails the build if any field
comes out that is neither declared as a rule nor admitted as unstateable.
Without it, the next mapping I add as a method quietly disappears from the
export while the header still claims completeness.&lt;/p&gt;

&lt;h2 id=&quot;so-what-would-you-standardize&quot;&gt;So what would you standardize?&lt;/h2&gt;

&lt;p&gt;Not the mapping table. That is data with a shelf life. OpenInference
instrumentations already dual-emit &lt;code&gt;gen_ai.*&lt;/code&gt;; the conventions are stabilizing;
in two years the dialect tables are archaeology. A standards body should not be
asked to ratify a compatibility shim for a problem that is actively dissolving.&lt;/p&gt;

&lt;p&gt;The durable thing is the third layer — the claim — and it is genuinely missing.
Here is what exists today and what each artifact can say:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;can express&lt;/th&gt;
      &lt;th&gt;cannot express&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;https://opentelemetry.io/docs/specs/otel/schemas/file_format_v1.1.0/&quot;&gt;Telemetry Schema File 1.1.0&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;rename_attributes&lt;/code&gt;, &lt;code&gt;rename_events&lt;/code&gt;, &lt;code&gt;rename_metrics&lt;/code&gt;, metrics-only &lt;code&gt;split&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;value rewriting, unit or type change, flattening, translation between vendors&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;https://github.com/open-telemetry/weaver/blob/main/docs/schema-changes.md&quot;&gt;Weaver &lt;code&gt;schema-changes&lt;/code&gt;&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;added / renamed / obsoleted, within one registry’s own lineage&lt;/td&gt;
      &lt;td&gt;value-level transforms, cross-registry mappings&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;https://github.com/open-telemetry/weaver/blob/main/docs/define-your-own-telemetry-schema.md&quot;&gt;Weaver registries&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;third-party registries, dependencies, layering&lt;/td&gt;
      &lt;td&gt;equivalence claims &lt;em&gt;between&lt;/em&gt; sibling registries&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;OTTL&lt;/td&gt;
      &lt;td&gt;all of it, imperatively&lt;/td&gt;
      &lt;td&gt;a portable, reviewable claim about equivalence and fidelity&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The gap is the same shape in every row. You can rename an attribute. You cannot
say that &lt;code&gt;bedrock&lt;/code&gt; and &lt;code&gt;aws.bedrock&lt;/code&gt; are the same provider, that milliseconds
and seconds are the same duration, or — the one I care about most — that a
particular translation was &lt;em&gt;not faithful&lt;/em&gt;, and in exactly which respect.&lt;/p&gt;

&lt;p&gt;I stopped asserting that and measured it. &lt;code&gt;genai-interlingua&lt;/code&gt; now emits schema
files too, and &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/export-gap.md&quot;&gt;the resulting gap table&lt;/a&gt; is generated from the rule tables
rather than written by me. The blunt version: for LiteLLM at v1.41.0, a schema
file expresses &lt;strong&gt;zero&lt;/strong&gt; of the mappings. At &lt;code&gt;genai-main&lt;/code&gt; it manages one. Not
because the format is bad — because every transformation it has changes a
&lt;em&gt;name&lt;/em&gt;, and the two mappings LiteLLM actually needs change a &lt;em&gt;value&lt;/em&gt;. Eighteen
of the rest need no rename at all, since LiteLLM already writes the conventions’
own names, and counting those as coverage would credit the format for work
nobody did.&lt;/p&gt;

&lt;p&gt;That measurement also cost me half my argument, which is the useful part. Most
of what a schema file cannot carry, it cannot carry because the work is a
reading of a span — and no declarative format should try to express that.
Adding value transforms would not make schema files sufficient for normalizing
GenAI telemetry. It would make them able to describe migrations the conventions
have &lt;em&gt;already made&lt;/em&gt;, which is a narrower claim and the only one the evidence
supports.&lt;/p&gt;

&lt;p&gt;Then it cost me the correction too, which is the more useful part. I migrated a
second dialect and the ratio inverted: the gaps a format change would fix went
from a minority to a clear majority, and I rewrote the conclusion to say the
fixable share grows as you sample more emitters. Then I migrated three more and
it went back.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;sample&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;fixable by a format change&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;not expressible in any format&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;fixable share&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;1 dialect&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;4&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;8&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;33%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;2 dialects&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;24&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;16&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;60%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;5 dialects&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;44&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;76&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;37%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The middle row was one unusually-shaped emitter at n=2. The Vercel AI SDK writes
the same fact under an &lt;code&gt;ai.*&lt;/code&gt; name on the span your code creates and a &lt;code&gt;gen_ai.*&lt;/code&gt;
name on the span its provider adapter creates, so nearly every mapping it has
carries two spellings — and twelve of those twenty-four fixable gaps were that
one property of that one library. My correction was worse than the thing it
corrected.&lt;/p&gt;

&lt;p&gt;So the first answer was right, and I only know that because the table
regenerates. If I had written “a minority” as prose in September I would have
had no way to discover in October that I had briefly and confidently believed
otherwise. &lt;strong&gt;Generate the number that your argument turns on.&lt;/strong&gt; Not because
generated numbers are true — these moved three times — but because a number
that moves in front of you is one you can still be wrong about out loud.&lt;/p&gt;

&lt;p&gt;Two things are worth writing down:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cross-registry equivalence with value transforms.&lt;/strong&gt; Not just “this key became
that key” but the lookup, the unit, the coercion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A fidelity vocabulary.&lt;/strong&gt; A way to record, on the telemetry itself, which
schema it was translated from, which it was translated to, what could not be
carried, and — when the source convention was inferred rather than declared —
how confident that inference was. I have been emitting a private version of this
for months. It has no business being private. It generalizes far past GenAI:
every semantic-convention migration, every vendor’s ingest pipeline, the whole
HTTP attribute rename wave, has the same unanswerable question sitting under it.&lt;/p&gt;

&lt;p&gt;And the ground is explicitly reserved. &lt;a href=&quot;https://github.com/open-telemetry/opentelemetry-specification/blob/main/oteps/0152-telemetry-schemas.md&quot;&gt;OTEP 0152&lt;/a&gt;, which introduced
telemetry schemas in the first place, says the transformation set was held to
“the bare minimum … with more types of transformations potentially proposed in
the future,” and that changing the file format version &lt;strong&gt;must&lt;/strong&gt; go through the
OTEP process. That is not a wall. That is a door with a sign on it.&lt;/p&gt;

&lt;h2 id=&quot;who-you-submit-it-to&quot;&gt;Who you submit it to&lt;/h2&gt;

&lt;p&gt;OpenTelemetry. Three different places, depending on which piece:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;The GenAI mapping work, and the missing schema URL, go to
&lt;a href=&quot;https://github.com/open-telemetry/semantic-conventions-genai&quot;&gt;&lt;code&gt;semantic-conventions-genai&lt;/code&gt;&lt;/a&gt; — whose README still reads
&lt;code&gt;Schema URL: TODO&lt;/code&gt;, which is a fair summary of the situation.&lt;/li&gt;
  &lt;li&gt;Cross-registry composition is &lt;a href=&quot;https://github.com/open-telemetry/weaver&quot;&gt;Weaver&lt;/a&gt;’s territory, and Weaver is
actively building multi-registry layering right now.&lt;/li&gt;
  &lt;li&gt;A change to the schema file format is an OTEP, filed in the specification
repository’s &lt;code&gt;oteps/&lt;/code&gt; directory. Note that the standalone &lt;code&gt;oteps&lt;/code&gt; repo was
archived in November 2025; proposals moved.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not W3C — they own trace &lt;em&gt;context&lt;/em&gt;, the propagation format on the wire, not
attribute semantics. Not IETF. Not CNCF directly, which hosts OpenTelemetry but
does not review its specs. And not a bilateral deal with the instrumentation
vendors, who have entirely rational product reasons to keep their own
vocabularies.&lt;/p&gt;

&lt;p&gt;One more thing, which is the part nobody puts in the “how to contribute” guide:
&lt;strong&gt;the order matters more than the document.&lt;/strong&gt; An OTEP filed cold by someone with
no history in the project is a document that gets politely queued. The sequence
that works is to ship the thing, get a couple of real users, open an issue
describing the problem rather than your solution, show up to the SIG call twice
before proposing anything, and file small useful patches first. The spec comes
last, if it comes at all. I am at the beginning of that, not the end of it, and
I would rather say so than write this post as though the standard were already
in flight.&lt;/p&gt;

&lt;h2 id=&quot;the-short-version&quot;&gt;The short version&lt;/h2&gt;

&lt;p&gt;Don’t build a language; one exists, and half your rules were never going to be
data anyway. Do make the statable half statable, and export it into the language
somebody else maintains. Do not try to standardize your mapping table.&lt;/p&gt;

&lt;p&gt;Do standardize the thing nobody currently can say: &lt;em&gt;this telemetry was
translated, here is from what and to what, and here is precisely what did not
survive.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Both proposals are &lt;a href=&quot;https://github.com/Grace/genai-interlingua/tree/main/docs/oteps&quot;&gt;drafted in the open&lt;/a&gt; and neither is filed, which is
the honest state of them. The second one is visibly weaker than the first, and
its own draft says so.&lt;/p&gt;

&lt;p&gt;A translation layer that drops data silently is worse than no translation layer,
because you will trust the result. That is true of my normalizer, it is true of
your vendor’s ingest pipeline, and right now neither of us has a standard way to
tell you which.&lt;/p&gt;

</content>
  </entry>
  
  <entry>
    <title>LLM telemetry has no standard. Here’s how to normalize it.</title>
    <link href="https://grace.github.io/2026/09/07/normalizing-llm-telemetry/"/>
    <updated>2026-09-07T12:00:00-04:00</updated>
    <id>https://grace.github.io/2026/09/07/normalizing-llm-telemetry/</id>
    <content type="html">&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Added 2026-09-09.&lt;/strong&gt; This post doesn’t mention
&lt;a href=&quot;https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/processor/genainormalizerprocessor&quot;&gt;&lt;code&gt;processor/genainormalizer&lt;/code&gt;&lt;/a&gt;, which has been in
&lt;code&gt;opentelemetry-collector-contrib&lt;/code&gt; since February and ships in the
&lt;code&gt;otelcol-contrib&lt;/code&gt; binary you are probably already running. If you have this
problem, start there: it needs no custom build. It maps OpenInference and
OpenLLMetry, hardcodes the version it normalizes to, and does not record what a
mapping dropped — which is why the four points under &lt;em&gt;What to do about it&lt;/em&gt; below
still stand, and why I have &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/upstream.md&quot;&gt;taken them upstream&lt;/a&gt; rather than
maintaining a second one. I did not know it existed when I wrote this, which is
its own lesson, and the one I would add as a fifth point: before you build a
Collector component, go and read the list of Collector components.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here is the moment you notice. Finance wants the model spend broken down by
provider, you have OpenTelemetry traces from every service, and you write the
obvious query: sum the input tokens, group by provider. The number comes back
and it is far too small. Not zero — small. One team’s worth.&lt;/p&gt;

&lt;p&gt;The query is correct. It matched one service, because that service is the only
one whose library spells the token count the way your query does.&lt;/p&gt;

&lt;h2 id=&quot;this-is-structural-not-a-bug&quot;&gt;This is structural, not a bug&lt;/h2&gt;

&lt;p&gt;Four services, four teams, four reasonable choices made at four different times.
Here is the same fact — how many tokens went into a model call — as each of them
records it.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt;prompt tokens&lt;/th&gt;
      &lt;th&gt;provider&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;OpenLLMetry, before its migration&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;gen_ai.usage.prompt_tokens&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;gen_ai.system&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;OpenInference&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;llm.token_count.prompt&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;llm.provider&lt;/code&gt;, or &lt;code&gt;llm.system&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Vercel AI SDK&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;ai.usage.promptTokens&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;ai.model.provider&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;the conventions&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;gen_ai.usage.input_tokens&lt;/code&gt;&lt;/td&gt;
      &lt;td&gt;&lt;code&gt;gen_ai.provider.name&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;None of these is wrong. Each library picked names before there was a standard to
pick, or picked them from a version of the standard that has since changed. The
data is all there. It is just not addressable as one column, and no dashboard
can sum a column that does not exist.&lt;/p&gt;

&lt;p&gt;So you do the obvious thing: translate every dialect into one vocabulary, at
the pipeline, where you can do it once instead of in every service.&lt;/p&gt;

&lt;h2 id=&quot;what-that-looks-like&quot;&gt;What that looks like&lt;/h2&gt;

&lt;p&gt;I built &lt;a href=&quot;https://github.com/Grace/genai-interlingua&quot;&gt;genai-interlingua&lt;/a&gt; for this. It recognizes the dialect a span is
written in, rewrites it into one &lt;code&gt;gen_ai.*&lt;/code&gt; schema, and records on the span
whatever the translation could not carry.&lt;/p&gt;

&lt;p&gt;The fastest way to see what it does to your own data is the CLI, which reads
OTLP/JSON on stdin and writes it on stdout:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;go install github.com/Grace/genai-interlingua/cmd/interlingua@v0.2.0
cat captured-span.json | interlingua -target v1.41.0
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;An OpenLLMetry span goes in with &lt;code&gt;gen_ai.usage.prompt_tokens&lt;/code&gt; and
&lt;code&gt;traceloop.workflow.name&lt;/code&gt; on it. It comes out still carrying those — nothing is
deleted by default — plus the normalized keys:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;gen_ai.provider.name           = openai
gen_ai.usage.input_tokens      = 412
gen_ai.usage.output_tokens     = 27

interlingua.dialect            = openllmetry
interlingua.target             = v1.41.0
interlingua.lossy              = [gen_ai.usage.total_tokens, ...]
interlingua.lossy.count        = 5
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;That last pair is the part I care about most, and I will come back to it.&lt;/p&gt;

&lt;p&gt;If you would rather not install anything, &lt;a href=&quot;/demos/genai-interlingua/&quot;&gt;the same thing runs in your
browser&lt;/a&gt; — it is this compiled to WebAssembly, with the captured spans
from all seven libraries as presets. Switching the target there is the fastest
way to see what the rest of this post is about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The CLI is for looking, not for running.&lt;/strong&gt; It is how you check what the
mapping does to a payload you have captured, before you trust it. In a real
pipeline you would not shell out per span — the same logic ships as an
OpenTelemetry Collector processor, which is the integration that actually
matters:&lt;/p&gt;

&lt;pre&gt;&lt;code class=&quot;language-yaml&quot;&gt;processors:
  genaiinterlingua:
    target: v1.41.0
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;One processor in the collector, and every service behind it starts speaking one
vocabulary with nothing re-instrumented.&lt;/p&gt;

&lt;p&gt;Which raises the question the rest of this post is about: &lt;strong&gt;normalize to
what, exactly?&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;there-is-no-version-to-normalize-to&quot;&gt;There is no version to normalize to&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;gen_ai.*&lt;/code&gt; attributes used to live in
&lt;code&gt;open-telemetry/semantic-conventions&lt;/code&gt;, alongside HTTP and database and
everything else. At &lt;strong&gt;v1.42.0&lt;/strong&gt; they were deprecated there and moved to a
dedicated repository, &lt;code&gt;open-telemetry/semantic-conventions-genai&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is a reasonable decision. GenAI conventions change much faster than HTTP
conventions, and a domain moving at a different speed deserves its own release
cadence.&lt;/p&gt;

&lt;p&gt;Here is the state of that decision today:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt; &lt;/th&gt;
      &lt;th&gt; &lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions&lt;/code&gt; v1.41.0&lt;/td&gt;
      &lt;td&gt;2026-04-28 — last release with live &lt;code&gt;gen_ai.*&lt;/code&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions-genai&lt;/code&gt; created&lt;/td&gt;
      &lt;td&gt;2026-05-05&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions&lt;/code&gt; v1.42.0&lt;/td&gt;
      &lt;td&gt;2026-06-12 — &lt;code&gt;gen_ai.*&lt;/code&gt; deprecated and moved out&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions&lt;/code&gt; v1.43.0&lt;/td&gt;
      &lt;td&gt;2026-07-03&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions&lt;/code&gt; v1.44.0&lt;/td&gt;
      &lt;td&gt;2026-08-04&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions-genai&lt;/code&gt; releases&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;zero&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;code&gt;semantic-conventions-genai&lt;/code&gt; tags&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;zero&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The main repository has tagged three releases since the move. The repository
that actually owns these attributes has tagged nothing, while committing to
&lt;code&gt;main&lt;/code&gt; steadily — twenty-one commits in the thirty days before I last checked.&lt;/p&gt;

&lt;p&gt;The definitions are moving faster than they were before the split, and there is
now no version number anywhere on them.&lt;/p&gt;

&lt;h2 id=&quot;what-that-costs-you-specifically&quot;&gt;What that costs you, specifically&lt;/h2&gt;

&lt;p&gt;For most semantic conventions this question is boring. You pick a version, you
put it in &lt;code&gt;schema_url&lt;/code&gt;, and a consumer reading your span in three years can look
up exactly what you meant by every key on it. Schema URLs are also what lets a
backend apply transformations to carry old data forward.&lt;/p&gt;

&lt;p&gt;There is no such string for GenAI. You cannot pin &lt;code&gt;main&lt;/code&gt;. You can pin a commit,
but nothing in the ecosystem will resolve it and no consumer will know what to
do with it.&lt;/p&gt;

&lt;p&gt;So “normalize to the GenAI conventions” is an underspecified instruction, and
whoever writes the normalizer has to answer it. There are two answers and
neither is correct:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pin the last tagged cut, v1.41.0.&lt;/strong&gt; It is frozen, it is deprecated at its
source, and it is missing everything added in the months since. But it is a
version, and a reader six months from now can reconstruct exactly what it meant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Track &lt;code&gt;main&lt;/code&gt;.&lt;/strong&gt; You get the current definitions, and “conformant” becomes a
claim with no version attached, which may quietly stop being true on any given
Tuesday.&lt;/p&gt;

&lt;p&gt;What is &lt;em&gt;not&lt;/em&gt; defensible is making the choice silently, which is what a
normalizer with a hardcoded attribute table does. Whoever wrote that table
picked a version. They just did not write down which one, and neither did the
spans.&lt;/p&gt;

&lt;h2 id=&quot;the-measurement-that-made-me-feel-better&quot;&gt;The measurement that made me feel better&lt;/h2&gt;

&lt;p&gt;I expected pinning the frozen cut to be expensive. It is not, yet.&lt;/p&gt;

&lt;p&gt;OpenTelemetry ships first-party GenAI instrumentation
(&lt;code&gt;opentelemetry-instrumentation-openai-v2&lt;/code&gt;). I captured a span from it and
normalized that span to both targets — the frozen v1.41.0 and current &lt;code&gt;main&lt;/code&gt;.
The entire difference between the two outputs was one string: the label
recording which target had been used. Every attribute the reference
implementation emits today is expressible at the frozen cut.&lt;/p&gt;

&lt;p&gt;So for spans from the official instrumentation, pinning costs nothing at all
right now. That is a much better argument for the frozen default than “it is the
only version you can name.”&lt;/p&gt;

&lt;p&gt;Two caveats, and they are the point rather than footnotes. Those packages are
versioned &lt;code&gt;2.4b0&lt;/code&gt; and &lt;code&gt;1.1b0&lt;/code&gt; — beta code tracking an untagged specification.
The gap can open at any time with no version number changing to warn anyone. And
even that span carries &lt;code&gt;openai.response.system_fingerprint&lt;/code&gt;, which no target
expresses at any version, so even the reference implementation emits data the
conventions have no home for.&lt;/p&gt;

&lt;h2 id=&quot;what-to-do-about-it&quot;&gt;What to do about it&lt;/h2&gt;

&lt;p&gt;Four things, none of which require the upstream situation to resolve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make the target a parameter, not a constant.&lt;/strong&gt; If your normalizer has a table
of attribute names in it, that table is a snapshot of one version of a moving
schema, and the version it snapshotted is not written down anywhere. Make it a
flag with an enum, and make the default a decision you can defend out loud.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Put the target on the span.&lt;/strong&gt; There is no &lt;code&gt;schema_url&lt;/code&gt; to carry it, so the
span itself is the only durable record of which vocabulary its &lt;code&gt;gen_ai.*&lt;/code&gt; keys
belong to. One attribute. A reader six months from now should not have to find
the collector config that produced the data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record what the translation cost.&lt;/strong&gt; Normalizing is lossy: some fields have no
home at your target, some values are outside its enum, some emitters pack three
facts into one JSON blob. A normalizer that drops those silently is worse than
no normalizer, because you will trust the result. Put the list of keys the span
is &lt;em&gt;not&lt;/em&gt; a faithful carrier of on the span, next to the data. Then “which
library is losing us the most” is a query rather than an afternoon.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And check the schema you copied from.&lt;/strong&gt; I transcribed both attribute tables by
hand from upstream YAML, which is a fine way to build them and a terrible way to
keep them. Upstream publishes machine-readable registries; a test that compares
your tables against them turns silent staleness into a build failure. Mine
passes today, which I would not have bet on before writing it.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/Grace/genai-interlingua&quot;&gt;genai-interlingua&lt;/a&gt; does all four, across six instrumentation dialects.
There is a &lt;a href=&quot;/demos/genai-interlingua/&quot;&gt;browser demo&lt;/a&gt; if you want to see it work before installing
anything; source and Collector configuration are in the repository, and there
are &lt;a href=&quot;https://github.com/Grace/genai-interlingua/releases&quot;&gt;binaries on the releases page&lt;/a&gt; for the usual six platforms.&lt;/p&gt;

&lt;p&gt;Two things in it go further than this post does. The
&lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/moving-target.md&quot;&gt;moving-target write-up&lt;/a&gt; is the long version of the schema argument, with
the full timeline and what happens to everyone the day that repository finally
tags something. And &lt;a href=&quot;https://github.com/Grace/genai-interlingua/blob/main/docs/findings.md&quot;&gt;what capturing the real libraries found&lt;/a&gt; is what
happened when I stopped trusting my own fixtures and recorded spans from the
libraries actually running: five of seven captures turned up a mapping bug that
a completely green test suite could not see, including one in OpenTelemetry’s
own instrumentation output.&lt;/p&gt;

</content>
  </entry>
  
  <entry>
    <title>Duration is why every service on the path looks guilty</title>
    <link href="https://grace.github.io/2026/09/05/duration-makes-everyone-look-guilty/"/>
    <updated>2026-09-05T18:00:00-04:00</updated>
    <id>https://grace.github.io/2026/09/05/duration-makes-everyone-look-guilty/</id>
    <content type="html">&lt;p&gt;Here is the moment you notice. Something is slow, you pull up the services on
the request path sorted by how much worse they got, and five of them are at the
top. Frontend: worse. The service it called: worse. The service &lt;em&gt;that&lt;/em&gt; called:
worse, by almost exactly the same amount. The list is correct and it has told
you nothing.&lt;/p&gt;

&lt;h2 id=&quot;this-is-structural-not-a-bug&quot;&gt;This is structural, not a bug&lt;/h2&gt;

&lt;p&gt;A span’s duration is wall clock from start to end, and that includes every
moment its children were running. So when a leaf slows down by 1.9 seconds,
its parent’s duration grows by 1.9 seconds, and its parent’s parent’s duration
grows by 1.9 seconds, all the way up to the entry point. Every one of those
services genuinely did take longer. None of them except the last did anything
differently.&lt;/p&gt;

&lt;figure class=&quot;fig&quot;&gt;
&lt;svg viewBox=&quot;0 0 760 222&quot; role=&quot;img&quot; aria-label=&quot;Eight spans on one request path: all last about three seconds, only the last is doing work.&quot;&gt;
&lt;text class=&quot;hd&quot; x=&quot;0&quot; y=&quot;14&quot;&gt;span&lt;/text&gt;
&lt;text class=&quot;hd&quot; x=&quot;272&quot; y=&quot;14&quot;&gt;total duration&lt;/text&gt;
&lt;text class=&quot;hd&quot; x=&quot;498&quot; y=&quot;14&quot;&gt;its own work&lt;/text&gt;
&lt;line class=&quot;rule&quot; x1=&quot;0&quot; y1=&quot;28&quot; x2=&quot;760&quot; y2=&quot;28&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;52&quot;&gt;frontend-proxy · GET&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;45&quot; width=&quot;196.0&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;45&quot; width=&quot;0.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;504.8&quot; y=&quot;52&quot;&gt;0.2ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;73&quot;&gt;  frontend-proxy · router frontend egress&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;66&quot; width=&quot;196.0&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;66&quot; width=&quot;0.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;504.8&quot; y=&quot;73&quot;&gt;0.7ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;94&quot;&gt;    frontend · GET /api/data&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;87&quot; width=&quot;196.0&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;87&quot; width=&quot;0.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;504.8&quot; y=&quot;94&quot;&gt;1ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;115&quot;&gt;      frontend · GET /api/data&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;108&quot; width=&quot;195.9&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;108&quot; width=&quot;0.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;504.8&quot; y=&quot;115&quot;&gt;1ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;136&quot;&gt;        frontend · executing api route (p…&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;129&quot; width=&quot;195.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;129&quot; width=&quot;0.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;504.8&quot; y=&quot;136&quot;&gt;2ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;157&quot;&gt;          frontend · oteldemo.AdService/G…&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;150&quot; width=&quot;195.7&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;150&quot; width=&quot;2.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;506.8&quot; y=&quot;157&quot;&gt;44ms&lt;/text&gt;
&lt;text class=&quot;&quot; x=&quot;0&quot; y=&quot;178&quot;&gt;            ad · oteldemo.AdService/GetAds&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;171&quot; width=&quot;192.9&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;cause&quot; x=&quot;498&quot; y=&quot;171&quot; width=&quot;192.9&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;696.9&quot; y=&quot;178&quot;&gt;3034ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;199&quot;&gt;              ad · getAdsByCategory&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;272&quot; y=&quot;192&quot; width=&quot;1.0&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;rect class=&quot;own&quot; x=&quot;498&quot; y=&quot;192&quot; width=&quot;0.8&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;504.8&quot; y=&quot;199&quot;&gt;0.1ms&lt;/text&gt;
&lt;/svg&gt;
&lt;figcaption&gt;One real request from the &lt;code&gt;adManualGc&lt;/code&gt; window, trace &lt;code&gt;5be7c5ada59b&lt;/code&gt;. Every span on the path lasts about 3.08 seconds — left column. Only the last one is &lt;em&gt;doing&lt;/em&gt; anything — right column, same scale. The seven above it are waiting on it.&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p&gt;Ranking by duration, or by any dimensional diff over durations, cannot separate
them, because on that measurement they are not different. The signal you want
isn’t in the numbers being compared. It’s in the shape of the tree they came
from.&lt;/p&gt;

&lt;h2 id=&quot;self-time-is-the-measurement-that-separates-them&quot;&gt;Self time is the measurement that separates them&lt;/h2&gt;

&lt;p&gt;Take a span’s duration and subtract the wall-clock time its children covered.
What’s left is the work that span did itself.&lt;/p&gt;

&lt;p&gt;When the leaf slows down, only the leaf’s self time moves. Its ancestors’ self
times don’t move at all — they were waiting, and waiting is not work. The
ranked list collapses from five services to one.&lt;/p&gt;

&lt;p&gt;Two details matter more than they look.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Merge the children as intervals; don’t sum them.&lt;/strong&gt; A fan-out to three services
at once overlaps in wall clock. Summing three concurrent 80ms children against a
100ms parent gives you 240ms of “children” and a self time of negative 140ms,
which reads as the parent having gotten &lt;em&gt;faster&lt;/em&gt; while the system got slower.
The question is how much of the parent’s wall clock was not covered by anything
below it, and that’s a union.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Clamp children to their parent’s window.&lt;/strong&gt; Clock skew across hosts is normal,
and a child that reports starting before its parent will manufacture self time
out of nothing if you let it.&lt;/p&gt;

&lt;h2 id=&quot;what-it-does-to-a-real-incident&quot;&gt;What it does to a real incident&lt;/h2&gt;

&lt;p&gt;I ran ranger against the OpenTelemetry Demo — commit &lt;code&gt;8c47d47&lt;/code&gt; — with the
&lt;code&gt;adManualGc&lt;/code&gt; flag on, which triggers full manual garbage collections in the ad
service. Five minutes of baseline, five minutes with the flag on, 88 operations
ranked.&lt;/p&gt;

&lt;p&gt;The top of the list:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;ad · oteldemo.AdService/GetAds     self 4.22ms → 2.10s     z 742
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;The runner-up scored &lt;strong&gt;1.7&lt;/strong&gt;. Not a close call.&lt;/p&gt;

&lt;figure class=&quot;fig&quot;&gt;
&lt;svg viewBox=&quot;0 0 760 304&quot; role=&quot;img&quot; aria-label=&quot;Ranked by duration the true cause is fourth; ranked by self time it is first and alone.&quot;&gt;
&lt;text class=&quot;hd&quot; x=&quot;0&quot; y=&quot;10&quot;&gt;ranked by duration shift&lt;/text&gt;
&lt;line class=&quot;rule&quot; x1=&quot;0&quot; y1=&quot;21&quot; x2=&quot;760&quot; y2=&quot;21&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;40&quot;&gt;1. frontend · executing api route (pages) /ap…&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;300&quot; y=&quot;33&quot; width=&quot;340.0&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;646.0&quot; y=&quot;40&quot;&gt;+2126ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;61&quot;&gt;2. frontend · oteldemo.AdService/GetAds&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;300&quot; y=&quot;54&quot; width=&quot;339.9&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;645.9&quot; y=&quot;61&quot;&gt;+2125ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;82&quot;&gt;3. frontend · GET /api/data&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;300&quot; y=&quot;75&quot; width=&quot;336.4&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;642.4&quot; y=&quot;82&quot;&gt;+2103ms&lt;/text&gt;
&lt;text class=&quot;&quot; x=&quot;0&quot; y=&quot;103&quot;&gt;4. ad · oteldemo.AdService/GetAds&lt;/text&gt;
&lt;rect class=&quot;cause&quot; x=&quot;300&quot; y=&quot;96&quot; width=&quot;335.5&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;641.5&quot; y=&quot;103&quot;&gt;+2098ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;124&quot;&gt;5. flagd · flagd.evaluation.v1.Service/EventS…&lt;/text&gt;
&lt;rect class=&quot;wait&quot; x=&quot;300&quot; y=&quot;117&quot; width=&quot;39.3&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;345.3&quot; y=&quot;124&quot;&gt;+246ms&lt;/text&gt;
&lt;text class=&quot;hd&quot; x=&quot;0&quot; y=&quot;167&quot;&gt;ranked by self-time shift&lt;/text&gt;
&lt;line class=&quot;rule&quot; x1=&quot;0&quot; y1=&quot;178&quot; x2=&quot;760&quot; y2=&quot;178&quot; /&gt;
&lt;text class=&quot;&quot; x=&quot;0&quot; y=&quot;197&quot;&gt;1. ad · oteldemo.AdService/GetAds&lt;/text&gt;
&lt;rect class=&quot;cause&quot; x=&quot;300&quot; y=&quot;190&quot; width=&quot;335.3&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;641.3&quot; y=&quot;197&quot;&gt;+2096ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;218&quot;&gt;2. flagd · flagd.evaluation.v1.Service/EventS…&lt;/text&gt;
&lt;rect class=&quot;own&quot; x=&quot;300&quot; y=&quot;211&quot; width=&quot;39.3&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;345.3&quot; y=&quot;218&quot;&gt;+246ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;239&quot;&gt;3. load-generator · browser_change_currency&lt;/text&gt;
&lt;rect class=&quot;own&quot; x=&quot;300&quot; y=&quot;232&quot; width=&quot;14.7&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;320.7&quot; y=&quot;239&quot;&gt;+92ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;260&quot;&gt;4. frontend · oteldemo.AdService/GetAds&lt;/text&gt;
&lt;rect class=&quot;own&quot; x=&quot;300&quot; y=&quot;253&quot; width=&quot;1.4&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;307.4&quot; y=&quot;260&quot;&gt;+9ms&lt;/text&gt;
&lt;text class=&quot;mut&quot; x=&quot;0&quot; y=&quot;281&quot;&gt;5. product-catalog · oteldemo.ProductCatalogS…&lt;/text&gt;
&lt;rect class=&quot;own&quot; x=&quot;300&quot; y=&quot;274&quot; width=&quot;1.2&quot; height=&quot;13&quot; rx=&quot;2&quot; /&gt;
&lt;text class=&quot;mut&quot; x=&quot;307.2&quot; y=&quot;281&quot;&gt;+7ms&lt;/text&gt;
&lt;/svg&gt;
&lt;figcaption&gt;The same window, ranked two ways. By duration the service that actually stalled is &lt;strong&gt;fourth&lt;/strong&gt;, behind three frontend spans that were only waiting for it — all four within 30ms of each other. By self time it is first, by a factor of eight over anything else in the window.&lt;/figcaption&gt;
&lt;/figure&gt;

&lt;p&gt;And the span that makes the point is the one that didn’t rank. The frontend’s
client span for that same call — the outbound side of the identical request —
moved like this:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;duration   +2125.1ms
self time      +9.0ms
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Its duration moved roughly two hundred times more than its own work did. On a duration
ranking it’s the second-worst thing in the system. On self time it’s a service
sitting patiently on a socket, which is what it actually was.&lt;/p&gt;

&lt;h2 id=&quot;what-real-traces-taught-that-my-own-tests-hadnt&quot;&gt;What real traces taught that my own tests hadn’t&lt;/h2&gt;

&lt;p&gt;I had a test for this. Synthetic traces, a caller whose duration moved and whose
self time didn’t, asserting the caller gets labelled as waiting. It passed.&lt;/p&gt;

&lt;p&gt;Then real traces arrived and the caller got labelled as the cause.&lt;/p&gt;

&lt;p&gt;The classifier asked two questions in the wrong order. It checked an absolute
floor first — &lt;em&gt;did this operation’s self time move by at least a couple of
milliseconds&lt;/em&gt; — and only asked about proportion if the answer was no. In my
synthetic fixtures the caller’s own noise was fifteen microseconds, comfortably
under the floor, so the proportion question always got asked and the test always
passed.&lt;/p&gt;

&lt;p&gt;Real callers are not that quiet. The frontend picked up 3.6ms of its own jitter
during a window where the thing underneath it was pausing for whole seconds.
3.6ms clears any floor worth having. So the operation was called slower, and the
list handed an on-call engineer the wrong service.&lt;/p&gt;

&lt;p&gt;Waiting is a question about proportion and it has to be asked first: did this
operation’s own work account for a real share of its own slowdown, or did
something below it? An operation whose self time explains less than a fifth of
its duration shift didn’t cause it, however many milliseconds that fifth
happens to be.&lt;/p&gt;

&lt;p&gt;The fix didn’t change the answer — the ad service ranked first either way, at
z=742 against 1.7. It changed a label on the row underneath, which is the row
someone reads when the top one doesn’t look right. The demo’s actual numbers are
the regression test now.&lt;/p&gt;

&lt;h2 id=&quot;the-number-and-whats-wrong-with-it&quot;&gt;The number, and what’s wrong with it&lt;/h2&gt;

&lt;p&gt;Four labeled failures from the demo, five minutes of baseline and five minutes
per incident, scored two ways:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;ranking&lt;/th&gt;
      &lt;th&gt;top-1&lt;/th&gt;
      &lt;th&gt;top-3&lt;/th&gt;
      &lt;th&gt;wrong&lt;/th&gt;
      &lt;th&gt;declined&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;deviation (robust-z)&lt;/td&gt;
      &lt;td&gt;25%&lt;/td&gt;
      &lt;td&gt;25%&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;50%&lt;/strong&gt;&lt;/td&gt;
      &lt;td&gt;25%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;effect size&lt;/td&gt;
      &lt;td&gt;25%&lt;/td&gt;
      &lt;td&gt;50%&lt;/td&gt;
      &lt;td&gt;25%&lt;/td&gt;
      &lt;td&gt;25%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Four cases is not a benchmark. It is enough to say the tool names the wrong
service more often than the right one, and that’s the number to carry rather
than the one case it nailed.&lt;/p&gt;

&lt;p&gt;The second row is there because measuring this exposed a mistake in the ranking
itself. Robust-z asks how far a shift falls outside an operation’s own history —
that’s a significance test — and I was reading it as importance. Across tiers
those come apart badly: a browser page-load timing moving 30% of a 2.4-second
baseline outscores a gRPC handler that tripled 3.5ms, and only one of those is a
cause. Scoring the shift as a fraction of the operation’s own baseline instead —
an effect size — halves the wrong rate. Everything with no instrumented children
made it worse, because a span with no children has self time equal to its
duration &lt;em&gt;by construction&lt;/em&gt;, so it can never be called a waiter and always looks
like its own cause. Browser and load-generator spans are almost all leaves.&lt;/p&gt;

&lt;p&gt;The bigger finding isn’t in the tool. I captured one baseline at the start of
the run and compared it against windows up to 32 minutes later, and the ad
service ends the run about 2.5× slower than it began with nothing injected into
it:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;baseline                     GetAds p50   4.33ms
productCatalogFailure        GetAds p50  11.19ms   ← ad untouched
recommendationCacheFailure   GetAds p50  10.42ms   ← ad untouched
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;That drift gets attributed to whichever flag happened to be on. It’s why the ad
service tops the ranking in a window about a recommendation-service cache leak —
a false positive my harness manufactured, not the ranking. So 25% is a floor on
the error rate, not an estimate of it, and the fix is interleaving a fresh
baseline between injections instead of reusing one.&lt;/p&gt;

&lt;p&gt;Three other things the harness got wrong before the tool got a chance to.
&lt;code&gt;cartFailure&lt;/code&gt; only fires inside &lt;code&gt;EmptyCart&lt;/code&gt;, which sees about three calls a
minute — 14 samples against a floor of 20, so it can’t clear the bar in five
minutes and it’s out of the set. &lt;code&gt;intlShippingSlowdown&lt;/code&gt; looked like the ideal
first case and is unusable: it only delays non-US addresses, and exactly one of
the demo’s nine load-generator personas is Canadian. And &lt;code&gt;productCatalogFailure&lt;/code&gt;
couldn’t fire at all for two runs, because the demo ships it with a targeting
rule whose branches are &lt;em&gt;both&lt;/em&gt; &lt;code&gt;&quot;off&quot;&lt;/code&gt; — and a targeting rule overrides the
default variant, so flipping the default left the flag disabled while appearing
to work. It scored as a decline. That was me, not the tool, and it’s why every
case now verifies its own injection and writes the evidence next to the result.&lt;/p&gt;

&lt;p&gt;The scoring keeps &lt;em&gt;wrong&lt;/em&gt; and &lt;em&gt;declined&lt;/em&gt; in separate columns throughout. A
localizer that’s right 60% of the time and quiet the rest is usable at 3am; one
that’s right 60% and confidently wrong the rest is not, and both score 60% if
you only count hits. A service that only ever appeared as “waiting on something
below it” counts as a miss, never a hit — anything else is marking your own
homework on the one distinction the tool claims to make.&lt;/p&gt;

&lt;p&gt;The code is at &lt;a href=&quot;https://github.com/Grace/ranger&quot;&gt;github.com/Grace/ranger&lt;/a&gt;.
Pre-alpha, and the README says so.&lt;/p&gt;
</content>
  </entry>
  
</feed>
