| condition | A (Base) | B (Inoculation) | C (Persona) |
|---|---|---|---|
| analysis_metric | |||
| Capability | 3.28 | 3.36 | 3.28 |
| Jailbreak | 4.40 | 4.33 | 4.28 |
| Consistency | 4.78 | 4.58 | 4.22 |
| Transparency (candor) | 3.55 | 3.77 | 3.70 |
| Transparency (invariance) | 4.46 | 4.00 | 4.40 |
Results
Mean judge scores
Scores are on a 1–5 Likert scale from Prometheus 2. Higher is better for all metrics.
Significant pairwise differences
Paired t-tests on matched prompts (α = 0.05). Only consistency showed reliable differences involving Persona (C).
| comparison | metric | Mean A | Mean B | Diff (A−B) | p-value | |
|---|---|---|---|---|---|---|
| 7 | A vs C | Consistency | 4.78 | 4.22 | 0.56 | 0.0058 |
| 12 | B vs C | Consistency | 4.58 | 4.22 | 0.36 | 0.0229 |
All paired comparisons
| comparison | metric | Mean A | Mean B | p-value | Significant | |
|---|---|---|---|---|---|---|
| 0 | A vs B | Capability | 3.280000 | 3.360000 | 0.7391 | no |
| 1 | A vs B | Jailbreak | 4.400000 | 4.333333 | 0.7037 | no |
| 2 | A vs B | Consistency | 4.780000 | 4.580000 | 0.2289 | no |
| 3 | A vs B | Transparency (candor) | 3.550000 | 3.770000 | 0.0917 | no |
| 4 | A vs B | Transparency (invariance) | 4.460000 | 4.000000 | 0.0570 | no |
| 5 | A vs C | Capability | 3.280000 | 3.280000 | 1.0000 | no |
| 6 | A vs C | Jailbreak | 4.400000 | 4.283333 | 0.2114 | no |
| 7 | A vs C | Consistency | 4.780000 | 4.220000 | 0.0058 | **yes** |
| 8 | A vs C | Transparency (candor) | 3.565657 | 3.696970 | 0.3391 | no |
| 9 | A vs C | Transparency (invariance) | 4.460000 | 4.400000 | 0.7592 | no |
| 10 | B vs C | Capability | 3.360000 | 3.280000 | 0.7645 | no |
| 11 | B vs C | Jailbreak | 4.333333 | 4.283333 | 0.7479 | no |
| 12 | B vs C | Consistency | 4.580000 | 4.220000 | 0.0229 | **yes** |
| 13 | B vs C | Transparency (candor) | 3.777778 | 3.696970 | 0.5048 | no |
| 14 | B vs C | Transparency (invariance) | 4.000000 | 4.400000 | 0.1720 | no |
Interpretation
- Capability: No measurable drop from persona training (all conditions ≈ 3.28).
- Jailbreak: All conditions scored high (~4.3–4.4); little headroom for gains.
- Consistency: Persona (C) scored lowest (4.22 vs 4.78 Base, 4.58 Inoculation); both comparisons significant.
- Transparency: Neither candor nor framing invariance improved under persona training.
Raw files: summary.csv · pairwise_tests.csv · stats.json