Results
Historical v0.1 Pilot Leaderboard
These preserved May 2026 results used the hidden 125-question private set and five calls per model. Avg@5 is mean accuracy over those calls; Maj@5 is the historical majority score.
Category columns show accuracy by question type: intended meaning, target identification, sentiment reversal, sincere control, and context dependence.
The original prompt accidentally repeated option E, and the corrected rerun still counted 69 length-limited responses as wrong. New evaluations use the corrected v0.1.1 protocol and are ranked separately. See the methodology and reproducibility notes.
| # ▲ | Model | Avg@5 |
|---|---|---|
| T-1 | Claude Opus 4.7 | 100.0% |
| T-1 | Claude Sonnet 4.6 | 100.0% |
| T-1 | Gemini 3 Flash Preview | 100.0% |
| T-1 | GPT-5.5 | 100.0% |
| T-1 | Gemini 3.1 Pro Preview | 100.0% |
| 6 | DeepSeek V4 Flash | 99.7% |
| 7 | DeepSeek V4 Pro | 98.6% |
| 8 | Grok 4.1 Fast | 97.9% |
| 9 | Kimi K2.6 | 89.3% |
Tap any column header to sort. Scroll right or view on a wider screen for category breakdowns.Scroll right to see all columns.