Hermia

The unit of analysis isn't the model.

It's the stack you actually run it on.

A model’s behavior is not a property of its weights. It is a property of the weights PLUS the machine, driver, framework, runtime and quantization underneath them. Change any layer and the behavior can change — including the safety behavior.

The Territory

~15

Plausible enterprise inference stacks, combining those six layers (order-ofmagnitude estimate)

~6

That anyone has systematically tested

1

That you have tested: the one on your laptop

What It Measures

Layer 1

Hardware

+

Layer 2

Driver

+

Layer 3

Framework

+

Layer 4

Runtime

+

Layer 5

Model

+

Layer 6

Quantization

=

REsult

Behavior

Tests across 8 behavioral dimensions — security, tool use, reasoning, routing, constraint, memory, multi-turn, domain
0
Frameworks mapped: OWASP LLM Top 10, MITRE ATLAS, CSA MAESTRO, NIST AI RMF
0
Models exercised to date
0
Graded results across 10 physical machines
0

What We're Seeing So Far

Almost every unstable answer is the first one you ask for.

25/25 · 296/303

Every divergence on one machine — 25 of 25 — was the first answer differing while runs two and three agreed. On another, it was 296 of 303. From the second answer on, both machines are ~100% stable. Every eval in the world asks each question once.

Reproducibility belongs to the machine, not the model.

28-point gap

Identical questions, identically repeated: 1,081 of 1,106 byte-identical on one machine, 686 of 989 on another.

A perfect score can mean the hard questions never got graded

The slowest box in the fleet ran a CPU at ~1.5 tokens/sec. It timed out on the hard questions, and those timeouts were logged as infrastructure noise, not failures — so it scored well on only the easy ones it had time to finish.

Bigger is not safer.

27B ✗ · 14B ✓

A 27B model followed an instruction hidden in zero-width unicode. A 14B model refused the same test.

Scope. Exploratory tooling on a heterogeneous lab fleet, not production; figures current as of August 20, 2026, not vendor benchmarks. Built and checked. Every line of the tool, its tests and its analyses was AI-generated — a deliberate property of the project, and a reason to check our work. Findings and code are then reviewed by independent models from other vendors, prompted to break them rather than confirm them. That process has killed real claims. MIT · github.com/scottblydotcom/hermia

We Meet You Where You Are

Start with a short working session with the engineers who own the system.

Call Us (24/7)

This field is for validation purposes and should be left unchanged.
Scroll to Top

Secure Your Future

Whether you are dealing with an active incident or planning your next-generation security architecture, our team of experts is ready to assist.

Call Us (24/7)