Engineering note · ongoing

Ask It Again

We checked how repeatable Jev is, built a model to explain what moved the answer, found the model was wrong, and came away with one check worth adding to anything built on it.

This is a running note, not a finished study — we're working out where Jev is useful and where it isn't, and we expect to keep adding to it. The one conclusion we'd act on so far: ask a second question about whether the text can be judged at all, because the confidence score cannot tell you that. Dated September 2026; the repository has the current state.3

Where this started

We read an article on dev.to about putting Jev behind a TLA+ specification and running 1,680 pharmacy decisions through it under chaos testing.2 TLA+ is a language for writing down precisely how a system is allowed to behave, so a checker can walk every reachable state and find the ones that break your rules.

One line in it stuck with us: the same request, sent five times, came back 0.03, 0.03, 0.03, 0.04, 0.04.

So the model doesn't always give the same answer to the same question. We wanted to know how much that matters for something we'd already shipped.

What Jev is

Jev, from TypeSafe, doesn't write text. It belongs to a different category — judgment models, or System One models — and gives back a number instead of prose.1

Ask it does this customer sound angry? and you get 0.9. Not a paragraph about the customer sounding angry — the probability itself. It answers in three shapes: a yes/no probability, a choice among options you define, or a score across ordered levels. That makes it something code can use directly, where software branches on a judgment: route this, escalate that, hold this one back.

What we did

The question we studied came from jev-voice, an extension we'd already built on Jev. It profiles writing along six dimensions — formality, authority, complexity, warmth, pace and conviction — and for each returns the full spread across the levels rather than a single label, so a text sitting 60% formal and 40% conversational is reported as that instead of rounded to a verdict.

We took one of those six — how formal is this writing?, on a four-level scale — and wrote thirty short texts. Eight obviously casual, eight obviously formal, fourteen deliberately in between.

Each text went through eight versions of the same request, twice each. Four versions changed something that cannot change the right answer: trailing whitespace, an extra field nothing reads, the fields in a different order. The other four reworded the question while keeping its meaning. Alongside that we ran a yes/no gate question on eight messages, and twenty texts where half weren't writing at all — a bare number, a URL, a base64 blob.

1,252 calls, one pinned model version so an update mid-study couldn't be mistaken for noise.

What we replicated

Most of what we measured was already public, and came out the same way.

Byte-identical requests don't always give byte-identical answers — but only in the middle. Texts at either end of the scale returned the same answer twenty times out of twenty. The in-between ones returned up to sixteen different answers in twenty calls.

And the three things that can differ between two calls come out in a consistent order:

What changed between callsmovement (sd, 0–3 scale)
nothing0.0102
something that can't matter0.0204
the wording0.0347

Published numbers for the same model version are ordered the same way, at a ratio of 1.74 where we got 1.70.2 That's a replication, not a discovery, and we'd rather say so than dress it up.

Two things we didn't expect

An unused field isn't free. Adding a field to the request that nothing reads moves the answer about twice as much as simply sending the same request again. Nothing in a request is inert, including the parts that aren't part of the question.

A lot of texts get exactly the same number. Of thirty texts, thirteen came back with identical values — six at the bottom of the scale, seven at the top. Not close; equal. Sort that list and the order inside those two blocks is invented. Allowing for movement, 17% of all pairs were too close to separate.

A four-level score carries less resolution than its decimals suggest, and anything mining those decimals is reading detail that isn't there. One text came back with confidence between 0.46 and 0.53 across identical requests, so a gate written confidence >= 0.5 lands on both sides of it between runs.

The model we built, and why we dropped it

In the raw data the causes are tangled together, so we fitted a model to separate them: treat every answer as the sum of a few hidden contributions — one for this text, one for this wording, one for this call — and work backwards from 1,252 answers to size each. We had it report each contribution as a range rather than a single number, which turned out to be the only part that mattered.

It said rewording moved the answer more than the changes that can't matter, and we were ready to write that down. Then we checked it against itself — a fitted model can generate fake data, and if the fake data doesn't resemble the real data, something is wrong. Ours produced data more spread out than what we'd actually measured.

The culprit was an assumption we hadn't noticed making: that every text reacts to rewording by about the same amount. Two of the fourteen in-between texts move by around 0.37 when reworded; the other twelve move by about 0.09. Forced into one number, the model had split the difference.

Refitting to allow a few texts to be much more sensitive than the rest fixed the check. And under that version, the range for the rewording effect reaches down to zero — so the finding we'd been about to publish doesn't survive its own follow-up. Had we reported a single number instead of a range, we would never have seen it.

AssumptionFits the dataRewording effect
all texts equally sensitiveno0.0022 [0.0007, 0.0041]
a few far more sensitiveyes0.0004 [0.0000, 0.0011]

What we were left with is that sensitivity to rewording belongs to particular questions rather than to the model — which, written down plainly, is close to obvious. Some texts are borderline and some aren't; the borderline ones move. We spent a fair amount of modelling to arrive at something you'd have guessed, and it rests on two texts out of fourteen. We're reporting it, not leaning on it.

The one thing worth taking away

We gave it a text and asked how formal the writing was. The text was the word yes. That's all — three letters, no sentence, nothing with a register.

It came back 98% confident: maximally casual writing.

Read generously the answer isn't even wrong — "yes" is not formal. But there was nothing there to be formal or informal about, and nothing in the response says so. This is the failure you can't watch for, because it doesn't look like uncertainty. A wobbling number tells you the model is unsure. This number is rock steady: ask again and you get 98% again. The model isn't hesitating. It's confidently answering a question that doesn't apply.

Put that in a pipeline and it stops being a curiosity. Say you're triaging email and scoring each one for tone before routing it. A reply arrives that reads, in full, yes. It gets a formality score like everything else — a real number, high confidence, no flag — and that number goes into a queue position, a priority, a routing rule. Nowhere in the chain is there anything saying there was no writing in this email. The pipeline can't tell the difference between a score that means something and a score that doesn't, because both arrive as a confident number in the same field.

Statistics has a name for this shape of mistake: an error of the third kind — the right answer to the wrong question. Measurement theory frames it as the difference between reliability and validity: whether a measurement repeats, versus whether it measures the thing you think it measures. Those come apart badly here. On the word yes the formality score is perfectly reliable — it returns 98% every time — and completely invalid.

So we added a second question to the same request. Not a threshold on the first one — a different question, asked at the same time, about the same text:

question:  Rate the formality of this writing. Consider word choice,
           sentence structure, contractions, colloquialisms, and register...

answerable: Does this text contain enough connected writing to support a
            judgement about its formality? Answer yes only if there is
            genuine prose present — sentences with word choice, structure
            and register that can actually be assessed. Answer no if the
            text is too short to have a register, is not prose at all, or
            contains no writing to assess.

  yes:  There is real prose here — at least one full sentence whose word
        choice, structure and register can be assessed.
  no:   There is nothing to assess: a bare value, an identifier or URL, a
        single word or fragment, punctuation, encoded data, or a bare list
        with no sentence.

On the word yes, that second question returns 2%.

We ran it over twenty texts: ten ordinary pieces of prose, and ten things that aren't writing at all — a bare number, a URL, TODO, a run of punctuation, a base64 blob, a list of constants. The second question separated them completely. Every piece of prose scored above 0.93; everything that wasn't writing scored below 0.31. Confidence didn't come close, and its mistakes went both ways:

TextConfidenceAnswerable
yes0.980.02
a customer-service reply0.500.98

The first row is the dangerous one — high confidence on something unjudgeable, which a gate waves straight through. The second is the expensive one: ordinary prose a confidence gate would escalate for no reason. Left unasked, the non-writing didn't get refused, it got rated: a list of mathematical constants came back at 2.40, more formal than a release note.

Writing one for your own question

The phrasing has to be specific to what you're asking, because "enough to answer" means something different each time. The pattern that worked for us: name what counts as enough, then name the kinds of thing that count as nothing. Here is the same check written for a completely different question — a yes/no gate on whether a message needs a human reply today:

question:  Does this message require a response from a human today?
           Answer yes only if delay until tomorrow would cause a real
           problem for the sender or for the business.

answerable: Does this text contain an actual message from someone, with
            enough content to judge whether it needs a reply?

  yes:  There is a real message here with enough content to assess.
  no:   There is nothing to assess: a bare value, an identifier, a
        fragment, or encoded data.

Both questions go in the same request, so they see exactly the same text and cost one call rather than two. Then the answer you act on is the first number, gated on the second — and when the second is low you have something specific to do, which is to stop rather than to guess.

The article we started from reaches the same wall and says so outright: a stability gate catches jitter around a value, but it cannot detect that no value is warranted — that needs a separate question, which they hadn't built.2 This is that question. It costs one extra question in a request you were already making. If you take one thing from this note, take that.

What the extension does

We packaged these checks as a swamp extension, so the same tests can be pointed at the next question instead of rebuilt each time.

You hand it a question and a set of example texts. It sends each text through the variants described above — the same request repeated, versions with something changed that shouldn't matter, versions with the wording changed — and, if you give it one, the second answerability question alongside. Then it reports back:

swamp extension pull @vcjdeboer/jev-reliability

swamp model method run @vcjdeboer/jev-reliability check myquestion \
  --input-file myquestion.json

swamp report get @vcjdeboer/jev-reliability-report --model myquestion

Two things it does deliberately. It hashes every request it sends and refuses to count two that turn out identical — which caught a mistake of ours, where a variant meant to reorder two fields left them in the order they were already in, so for 480 calls it was quietly just sending the same request again. We'd read that code several times without seeing it. And if the model fit hasn't settled, it says so instead of printing the numbers.

None of it says whether Jev is right — only whether it's consistent. That's a different question and it needs labelled examples.

Code, raw data and the full write-up are in the repository.3

References

  1. TypeSafe. Jev — System One models. docs.typesafe.ai ↩
  2. copyleftdev (2026). I Put Jev Behind a TLA+ Spec and Ran 1,680 Chaos-Tested Pharmacy Decisions. Zero Wrong Verdicts. dev.to ↩
  3. jev-reliability — extension, data and write-up. github.com/vcjdeboer/jev-reliability ↩
← vcjdeboer.github.io