Results

Historical v0.1 Pilot Leaderboard

These preserved May 2026 results used the hidden 125-question private set and five calls per model. Avg@5 is mean accuracy over those calls; Maj@5 is the historical majority score.

Category columns show accuracy by question type: intended meaning, target identification, sentiment reversal, sincere control, and context dependence.

The original prompt accidentally repeated option E, and the corrected rerun still counted 69 length-limited responses as wrong. New evaluations use the corrected v0.1.1 protocol and are ranked separately. See the methodology and reproducibility notes.

#ModelAvg@5
T-1Claude Opus 4.7100.0%
T-1Claude Sonnet 4.6100.0%
T-1Gemini 3 Flash Preview100.0%
T-1GPT-5.5100.0%
T-1Gemini 3.1 Pro Preview100.0%
6DeepSeek V4 Flash99.7%
7DeepSeek V4 Pro98.6%
8Grok 4.1 Fast97.9%
9Kimi K2.689.3%

Tap any column header to sort. Scroll right or view on a wider screen for category breakdowns.