HERMES IDENTITY / ERGONOMICS / CHEMISTRY / V4

EVALUATION INSTRUMENT

FOUR-CONDITION FACTORIAL COMPARISON

Two levers.
Four response surfaces.

Does Alfred’s SOUL change the answer? Does it change what the ADHD overlay adds? Same prompts, same model, both switches tested on and off.

“No custom SOUL” is not no identity. Hermes falls back to its built-in default identity. Machines do enjoy a technicality.

MODEL PROVIDER CASES

Condition matrix

ACCURACY / EFFICIENCY / WORDS
DESCRIPTIVE INTERACTION

What changed when both switches moved?

Calculating the interaction…

FOUR CONTROLLED CONTRASTS

Which lever won each case?

A winner needs two of three strict blind-judge verdicts. Ambiguity stays a tie rather than being massaged into a headline.

SPECIMEN SELECTOR

Evaluation cases

01 /

User prompt

Loading evaluation…
DIRECT CELL COMPARISON

Pick any two scenarios.

Identity and presentation mode can move independently. As nature intended.

CONDITION 00

Scenario A

SELECTED PAIR · WHOLE SUITE

How does this matchup perform overall?

to navigate

METHOD / MATCHED CONDITIONS

One model, controlled instruction layers.

Matched cellsCustom Alfred SOUL or Hermes’ built-in identity crossed with presentation modes.

Held constantModel, provider, memory, user profile, skills, tools, config, working directory, and prompts.

Gate firstThree blinded judge passes score every answer independently before the focal quality metric.

Directional evidenceOne generated sample per cell and same-model judges. Useful for mechanism hunting, not statistical scripture.