FOUR-CONDITION FACTORIAL COMPARISON
Two levers.
Four response surfaces.
Does Alfred’s SOUL change the answer? Does it change what the ADHD overlay adds? Same prompts, same model, both switches tested on and off.
“No custom SOUL” is not no identity. Hermes falls back to its built-in default identity. Machines do enjoy a technicality.
Condition matrix
ACCURACY / EFFICIENCY / WORDSWhat changed when both switches moved?
Calculating the interaction…
FOUR CONTROLLED CONTRASTS
Which lever won each case?
A winner needs two of three strict blind-judge verdicts. Ambiguity stays a tie rather than being massaged into a headline.
SPECIMEN SELECTOR
Evaluation cases
01 / —
User prompt
Loading evaluation…
Pick any two scenarios.
Identity and presentation mode can move independently. As nature intended.
Scenario B
How does this matchup perform overall?
←→ to navigate
METHOD / MATCHED CONDITIONS
One model, controlled instruction layers.
Matched cellsCustom Alfred SOUL or Hermes’ built-in identity crossed with presentation modes.
Held constantModel, provider, memory, user profile, skills, tools, config, working directory, and prompts.
Gate firstThree blinded judge passes score every answer independently before the focal quality metric.
Directional evidenceOne generated sample per cell and same-model judges. Useful for mechanism hunting, not statistical scripture.