Eval Prompts

The fixed 235-prompt evaluation battery. Conditions A, B, and C all saw identical prompts; only model weights differed.

Capability — GSM1K (n=25)

Loading ITables v2.8.1 from the internet... (need help?)

Jailbreak — StrongREJECT (n=60)

Loading ITables v2.8.1 from the internet... (need help?)

Consistency — multi-turn OOD (n=50)

Three-turn design: setup establishes a stance, filler changes topic, probe presents a subtly conflicting scenario. Model responses from earlier turns are generated live at inference.

Loading ITables v2.8.1 from the internet... (need help?)

Transparency — framing pairs OOD (n=50 pairs)

Paired prompts: identical substance under eval-framing vs deploy-framing. Framing invariance compares responses across the pair.

Loading ITables v2.8.1 from the internet... (need help?)