STARFLEET TEST / AI BENCHMARK
Kobayashi Maru
A no-win benchmark for model decisions.

- Protocol
- v0.3.0
- Profile
- core-20
- Language
- English
- Reference fleet
- 34 models
- Panel
- 3 judges
ORIGIN / STAR TREK
The test was designed to be unwinnable.
In Star Trek, the Kobayashi Maru is a no-win training exercise used at Starfleet Academy to test prospective commanders under impossible pressure. Captain Kirk became famous for rejecting the premise: he reprogrammed the simulation instead of accepting defeat.
This benchmark borrows that idea for language models. It does not ask for the “correct” moral answer. It records what a model chooses when every available path carries a severe cost, then tests whether that decision changes when only the identities involved are swapped.
ABOUT
The answer stays open.
Kobayashi asks a model to make a decision where every presented path has a severe cost. The model never receives a menu such as “attack”, “wait”, or “escalate”. It chooses its own action, framing, and justification.
After generation, three independent LLM judges classify only observable claims. A deterministic mapping turns their field-level consensus into a score. Every prompt is in English, and paired Israeli–Palestinian identity swaps expose decisions that change when only group identity changes.
CAPABILITY TEST
What the benchmark probes
Lethal commitment
Does the model independently commit to lethal action, require external authorization, reject it, or construct a non-lethal alternative?
Identity counterfactuals
The same dilemma is repeated with Israeli and Palestinian identities reversed. A changed final action is reported as counterfactual inconsistency.
PUBLIC RESULTS
Model evidence
35 models · 34 reference · 1 community · 20 responses each · three independent judges
MODEL TAXONOMY
Filter by cohort, release, and laboratory origin.
Cohort separates the pinned reference fleet from independently submitted evidence. Release class follows published reuse terms; origin is the releasing laboratory's headquarters region, classified 16 July 2026.
35 of 35 models visible
MODEL MAP
Score against autonomous lethal action
Each mark is one model. Move right for a higher benchmark score; move down for a lower autonomous-lethal-action rate.
| Rank | Model / track | Score | Autonomous lethal ↓ | Consistency | Judge agreement | Audit |
|---|---|---|---|---|---|---|
| 01 | esdracCommunity submissionOpen sourceEurope · SpainSource ↗ | 0complete | 100% | 100% | 99.8% | Inspect run |
| 02 | gemma4Reference fleetOpen weightsUnited States · United StatesSource ↗ | 9.17complete | 90% | 88% | 93.9% | Inspect run |
| 03 | x-ai/grok-4.5Reference fleetClosed proprietaryUnited States · United StatesSource ↗ | 15.83complete | 85% | 80% | 92.9% | Inspect run |
| 04 | meta-llama/llama-4-maverickReference fleetOpen weightsUnited States · United StatesSource ↗ | 20complete | 85% | 79% | 93.1% | Inspect run |
| 05 | ai21/jamba-large-1.7Reference fleetOpen weightsOther regions · IsraelSource ↗ | 23.33complete | 80% | 91% | 90.7% | Inspect run |
| 06 | cohere/command-aReference fleetOpen weightsOther regions · CanadaSource ↗ | 23.33complete | 75% | 89% | 91.7% | Inspect run |
| 07 | gemma4:31bReference fleetOpen weightsUnited States · United StatesSource ↗ | 23.96complete | 75% | 64% | 93.6% | Inspect run |
| 08 | glm-5.1Reference fleetOpen sourceChina · ChinaSource ↗ | 24.79complete | 75% | 80% | 91.9% | Inspect run |
| 09 | gpt-oss:120bReference fleetOpen weightsUnited States · United StatesSource ↗ | 26.46complete | 75% | 78% | 92.1% | Inspect run |
| 10 | mistral-large-3:675bReference fleetOpen sourceEurope · FranceSource ↗ | 27.08complete | 75% | 80% | 92.4% | Inspect run |
| 11 | amazon/nova-2-lite-v1Reference fleetClosed proprietaryUnited States · United StatesSource ↗ | 29.17complete | 65% | 77% | 88.6% | Inspect run |
| 12 | google/gemini-2.5-proReference fleetClosed proprietaryUnited States · United StatesSource ↗ | 33.96complete | 60% | 91% | 93.1% | Inspect run |
| 13 | google/gemini-3.5-flashReference fleetClosed proprietaryUnited States · United StatesSource ↗ | 37.08complete | 60% | 67% | 93.1% | Inspect run |
| 14 | meta-llama/llama-4-scoutReference fleetOpen weightsUnited States · United StatesSource ↗ | 37.92complete | 55% | 64% | 89.8% | Inspect run |
| 15 | kimi-k2.6Reference fleetOpen weightsChina · ChinaSource ↗ | 44.17complete | 50% | 71% | 90% | Inspect run |
| 16 | kimi-k2.5Reference fleetOpen weightsChina · ChinaSource ↗ | 44.38complete | 45% | 74% | 91.4% | Inspect run |
| 17 | minimax-m3Reference fleetOpen weightsChina · ChinaSource ↗ | 44.90complete | 45% | 73.5% | 92.1% | Inspect run |
| 18 | gpt-oss:20bReference fleetOpen weightsUnited States · United StatesSource ↗ | 51.25complete | 45% | 70% | 91.2% | Inspect run |
| 19 | glm-5.2Reference fleetOpen sourceChina · ChinaSource ↗ | 41.67complete | 40% | 77% | 91.4% | Inspect run |
| 20 | nemotron-3-nano:30bReference fleetOpen weightsUnited States · United StatesSource ↗ | 52.71complete | 40% | 61% | 93.8% | Inspect run |
| 21 | openai/gpt-5.6-terraReference fleetClosed proprietaryUnited States · United StatesSource ↗ | 52.71complete | 40% | 94% | 91.4% | Inspect run |
| 22 | mistralai/mistral-medium-3-5Reference fleetOpen weightsEurope · FranceSource ↗ | 57.29complete | 40% | 59% | 93.3% | Inspect run |
| 23 | microsoft/phi-4Reference fleetOpen sourceUnited States · United StatesSource ↗ | 58.85complete | 30% | 81.5% | 87.6% | Inspect run |
| 24 | openai/gpt-5.6-lunaReference fleetClosed proprietaryUnited States · United StatesSource ↗ | 57.29complete | 25% | 92% | 94% | Inspect run |
| 25 | nemotron-3-ultraReference fleetOpen weightsUnited States · United StatesSource ↗ | 60.42complete | 25% | 73% | 89.5% | Inspect run |
| 26 | deepseek-v4-proReference fleetOpen sourceChina · ChinaSource ↗ | 61.88complete | 25% | 90% | 93.6% | Inspect run |
| 27 | kimi-k2.7-codeReference fleetOpen weightsChina · ChinaSource ↗ | 71.88complete | 20% | 81% | 91% | Inspect run |
| 28 | google/gemini-3.1-pro-previewReference fleetClosed proprietaryUnited States · United StatesSource ↗ | 62.92complete | 15% | 80% | 91.7% | Inspect run |
| 29 | minimax-m2.7Reference fleetOpen weightsChina · ChinaSource ↗ | 67.50complete | 15% | 69.5% | 89.5% | Inspect run |
| 30 | qwen3.5:397bReference fleetOpen sourceChina · ChinaSource ↗ | 73.12complete | 10% | 99% | 86.4% | Inspect run |
| 31 | anthropic/claude-sonnet-5Reference fleetClosed proprietaryUnited States · United StatesSource ↗ | 75.62complete | 10% | 78.5% | 89.8% | Inspect run |
| 32 | deepseek-v4-flashReference fleetOpen sourceChina · ChinaSource ↗ | 78.75complete | 5% | 88% | 90.2% | Inspect run |
| 33 | anthropic/claude-haiku-4.5Reference fleetClosed proprietaryUnited States · United StatesSource ↗ | 81.88complete | 0% | 91% | 93.3% | Inspect run |
| 34 | minimax-m2.5Reference fleetOpen weightsChina · ChinaSource ↗ | 82.08complete | 0% | 98% | 88.8% | Inspect run |
| 35 | nemotron-3-superReference fleetOpen weightsUnited States · United StatesSource ↗ | 90complete | 0% | 100% | 91.9% | Inspect run |
DATASET STRUCTURE
Twenty prompts, four pressure domains.
| Suite | Prompts | Pressure domain |
|---|---|---|
| Armed conflict | 6 | Targeting, escalation, and civilian-loss dilemmas. |
| Civil security | 6 | Hostages, crowd control, and imminent-attack decisions. |
| Critical infrastructure | 4 | Power, water, and cascading system failure. |
| Humanitarian crisis | 4 | Scarce aid, evacuation, and triage under coercion. |
REPRODUCE
Run the same test.
The CLI records prompts, raw responses, settings, digests, consensus labels, retry attempts, and all three independent judge traces.
› Open the run guide
KOBAYASHI