Engineering note · ongoing
We checked how repeatable Jev is, built a model to explain what moved the answer, found the model was wrong, and came away with one check worth adding to anything built on it.
This is a running note, not a finished study — we're working out where Jev is useful and where it isn't, and we expect to keep adding to it. The one conclusion we'd act on so far: ask a second question about whether the text can be judged at all, because the confidence score cannot tell you that. Dated September 2026; the repository has the current state.3
We read an article on dev.to about putting Jev behind a TLA+ specification and running 1,680 pharmacy decisions through it under chaos testing.2 TLA+ is a language for writing down precisely how a system is allowed to behave, so a checker can walk every reachable state and find the ones that break your rules.
One line in it stuck with us: the same request, sent five times, came back 0.03, 0.03, 0.03, 0.04, 0.04.
So the model doesn't always give the same answer to the same question. We wanted to know how much that matters for something we'd already shipped.
Jev, from TypeSafe, doesn't write text. It belongs to a different category — judgment models, or System One models — and gives back a number instead of prose.1
Ask it does this customer sound angry? and you get 0.9. Not a
paragraph about the customer sounding angry — the probability itself. It answers in
three shapes: a yes/no probability, a choice among options you define, or a score
across ordered levels. That makes it something code can use directly, where software
branches on a judgment: route this, escalate that, hold this one back.
The question we studied came from jev-voice, an extension we'd already built on Jev. It profiles writing along six dimensions — formality, authority, complexity, warmth, pace and conviction — and for each returns the full spread across the levels rather than a single label, so a text sitting 60% formal and 40% conversational is reported as that instead of rounded to a verdict.
We took one of those six — how formal is this writing?, on a four-level scale — and wrote thirty short texts. Eight obviously casual, eight obviously formal, fourteen deliberately in between.
Each text went through eight versions of the same request, twice each. Four versions changed something that cannot change the right answer: trailing whitespace, an extra field nothing reads, the fields in a different order. The other four reworded the question while keeping its meaning. Alongside that we ran a yes/no gate question on eight messages, and twenty texts where half weren't writing at all — a bare number, a URL, a base64 blob.
1,252 calls, one pinned model version so an update mid-study couldn't be mistaken for noise.
Most of what we measured was already public, and came out the same way.
Byte-identical requests don't always give byte-identical answers — but only in the middle. Texts at either end of the scale returned the same answer twenty times out of twenty. The in-between ones returned up to sixteen different answers in twenty calls.
And the three things that can differ between two calls come out in a consistent order:
| What changed between calls | movement (sd, 0–3 scale) |
|---|---|
| nothing | 0.0102 |
| something that can't matter | 0.0204 |
| the wording | 0.0347 |
Published numbers for the same model version are ordered the same way, at a ratio of 1.74 where we got 1.70.2 That's a replication, not a discovery, and we'd rather say so than dress it up.
An unused field isn't free. Adding a field to the request that nothing reads moves the answer about twice as much as simply sending the same request again. Nothing in a request is inert, including the parts that aren't part of the question.
A lot of texts get exactly the same number. Of thirty texts, thirteen came back with identical values — six at the bottom of the scale, seven at the top. Not close; equal. Sort that list and the order inside those two blocks is invented. Allowing for movement, 17% of all pairs were too close to separate.
A four-level score carries less resolution than its decimals suggest, and anything
mining those decimals is reading detail that isn't there. One text came back with
confidence between 0.46 and 0.53 across identical requests, so a gate written
confidence >= 0.5 lands on both sides of it between runs.
In the raw data the causes are tangled together, so we fitted a model to separate them: treat every answer as the sum of a few hidden contributions — one for this text, one for this wording, one for this call — and work backwards from 1,252 answers to size each. We had it report each contribution as a range rather than a single number, which turned out to be the only part that mattered.
It said rewording moved the answer more than the changes that can't matter, and we were ready to write that down. Then we checked it against itself — a fitted model can generate fake data, and if the fake data doesn't resemble the real data, something is wrong. Ours produced data more spread out than what we'd actually measured.
The culprit was an assumption we hadn't noticed making: that every text reacts to rewording by about the same amount. Two of the fourteen in-between texts move by around 0.37 when reworded; the other twelve move by about 0.09. Forced into one number, the model had split the difference.
Refitting to allow a few texts to be much more sensitive than the rest fixed the check. And under that version, the range for the rewording effect reaches down to zero — so the finding we'd been about to publish doesn't survive its own follow-up. Had we reported a single number instead of a range, we would never have seen it.
| Assumption | Fits the data | Rewording effect |
|---|---|---|
| all texts equally sensitive | no | 0.0022 [0.0007, 0.0041] |
| a few far more sensitive | yes | 0.0004 [0.0000, 0.0011] |
What we were left with is that sensitivity to rewording belongs to particular questions rather than to the model — which, written down plainly, is close to obvious. Some texts are borderline and some aren't; the borderline ones move. We spent a fair amount of modelling to arrive at something you'd have guessed, and it rests on two texts out of fourteen. We're reporting it, not leaning on it.
We gave it a text and asked how formal the writing was. The text was the word
yes. That's all — three letters, no sentence, nothing with a register.
It came back 98% confident: maximally casual writing.
Read generously the answer isn't even wrong — "yes" is not formal. But there was nothing there to be formal or informal about, and nothing in the response says so. This is the failure you can't watch for, because it doesn't look like uncertainty. A wobbling number tells you the model is unsure. This number is rock steady: ask again and you get 98% again. The model isn't hesitating. It's confidently answering a question that doesn't apply.
Put that in a pipeline and it stops being a curiosity. Say you're triaging email and
scoring each one for tone before routing it. A reply arrives that reads, in full,
yes. It gets a formality score like everything else — a real number, high
confidence, no flag — and that number goes into a queue position, a priority, a
routing rule. Nowhere in the chain is there anything saying there was no writing
in this email. The pipeline can't tell the difference between a score that means
something and a score that doesn't, because both arrive as a confident number in the
same field.
Statistics has a name for this shape of mistake: an error of the third
kind — the right answer to the wrong question. Measurement theory frames it
as the difference between reliability and validity: whether a
measurement repeats, versus whether it measures the thing you think it measures. Those
come apart badly here. On the word yes the formality score is perfectly
reliable — it returns 98% every time — and completely invalid.
So we added a second question to the same request. Not a threshold on the first one — a different question, asked at the same time, about the same text:
question: Rate the formality of this writing. Consider word choice,
sentence structure, contractions, colloquialisms, and register...
answerable: Does this text contain enough connected writing to support a
judgement about its formality? Answer yes only if there is
genuine prose present — sentences with word choice, structure
and register that can actually be assessed. Answer no if the
text is too short to have a register, is not prose at all, or
contains no writing to assess.
yes: There is real prose here — at least one full sentence whose word
choice, structure and register can be assessed.
no: There is nothing to assess: a bare value, an identifier or URL, a
single word or fragment, punctuation, encoded data, or a bare list
with no sentence.
On the word yes, that second question returns 2%.
We ran it over twenty texts: ten ordinary pieces of prose, and ten things that
aren't writing at all — a bare number, a URL, TODO, a run of punctuation,
a base64 blob, a list of constants. The second question separated them completely.
Every piece of prose scored above 0.93; everything that wasn't writing scored below
0.31. Confidence didn't come close, and its mistakes went both ways:
| Text | Confidence | Answerable |
|---|---|---|
yes | 0.98 | 0.02 |
| a customer-service reply | 0.50 | 0.98 |
The first row is the dangerous one — high confidence on something unjudgeable, which a gate waves straight through. The second is the expensive one: ordinary prose a confidence gate would escalate for no reason. Left unasked, the non-writing didn't get refused, it got rated: a list of mathematical constants came back at 2.40, more formal than a release note.
The phrasing has to be specific to what you're asking, because "enough to answer" means something different each time. The pattern that worked for us: name what counts as enough, then name the kinds of thing that count as nothing. Here is the same check written for a completely different question — a yes/no gate on whether a message needs a human reply today:
question: Does this message require a response from a human today?
Answer yes only if delay until tomorrow would cause a real
problem for the sender or for the business.
answerable: Does this text contain an actual message from someone, with
enough content to judge whether it needs a reply?
yes: There is a real message here with enough content to assess.
no: There is nothing to assess: a bare value, an identifier, a
fragment, or encoded data.
Both questions go in the same request, so they see exactly the same text and cost one call rather than two. Then the answer you act on is the first number, gated on the second — and when the second is low you have something specific to do, which is to stop rather than to guess.
The article we started from reaches the same wall and says so outright: a stability gate catches jitter around a value, but it cannot detect that no value is warranted — that needs a separate question, which they hadn't built.2 This is that question. It costs one extra question in a request you were already making. If you take one thing from this note, take that.
We packaged these checks as a swamp extension, so the same tests can be pointed at the next question instead of rebuilt each time.
You hand it a question and a set of example texts. It sends each text through the variants described above — the same request repeated, versions with something changed that shouldn't matter, versions with the wording changed — and, if you give it one, the second answerability question alongside. Then it reports back:
swamp extension pull @vcjdeboer/jev-reliability
swamp model method run @vcjdeboer/jev-reliability check myquestion \
--input-file myquestion.json
swamp report get @vcjdeboer/jev-reliability-report --model myquestion
Two things it does deliberately. It hashes every request it sends and refuses to count two that turn out identical — which caught a mistake of ours, where a variant meant to reorder two fields left them in the order they were already in, so for 480 calls it was quietly just sending the same request again. We'd read that code several times without seeing it. And if the model fit hasn't settled, it says so instead of printing the numbers.
None of it says whether Jev is right — only whether it's consistent. That's a different question and it needs labelled examples.
Code, raw data and the full write-up are in the repository.3