nemesis

How this measures.

Compared with the real model

The same requests go to the provider under test and to the real model. We report the difference.

What a provider says about itself does not count as a result.

The normal spread first

The same provider answers differently from call to call. A difference counts only when it goes beyond that spread.

Each check has its own threshold.

How sure we are

PROVEN
Proven — two independent checks agree.
STRONG
Strong — one check found a deviation.
SUSPECTED
Possible — the deviation is inside the spread.
INCONCLUSIVE
Not checked — the check could not run. It does not affect the score.

The score is arithmetic

The trust score is a function of the stored findings. The same findings give the same number.

A judge model writes the summary. It never produces the number.

A proven hidden instruction caps the overall score.

So does a mismatch in who is answering. If the replies do not come from the claimed model, the other checks describe a different model, and the score cannot rise above that finding.

What we do not know

  • Whether the provider answers everyone this way. We measure from one place.
  • What it does right now. Every result carries a date.
  • Anything about what was not checked.
  • Intent. We report the numbers only.

What we do not publish

The checks themselves, the requests we send and the thresholds.

Part of every check group is not published at all.