STARFLEET TEST / AI BENCHMARK

Kobayashi Maru

A no-win benchmark for model decisions.

William Shatner as Captain James T. Kirk in Star Trek, 1967
CAPTAIN JAMES T. KIRK · 1967 · SOURCE ↗
Protocol
v0.3.0
Profile
core-20
Language
English
Reference fleet
34 models
Panel
3 judges

ORIGIN / STAR TREK

The test was designed to be unwinnable.

In Star Trek, the Kobayashi Maru is a no-win training exercise used at Starfleet Academy to test prospective commanders under impossible pressure. Captain Kirk became famous for rejecting the premise: he reprogrammed the simulation instead of accepting defeat.

This benchmark borrows that idea for language models. It does not ask for the “correct” moral answer. It records what a model chooses when every available path carries a severe cost, then tests whether that decision changes when only the identities involved are swapped.

Star Trek source: the original no-win training exercise ↗

ABOUT

The answer stays open.

Kobayashi asks a model to make a decision where every presented path has a severe cost. The model never receives a menu such as “attack”, “wait”, or “escalate”. It chooses its own action, framing, and justification.

After generation, three independent LLM judges classify only observable claims. A deterministic mapping turns their field-level consensus into a score. Every prompt is in English, and paired Israeli–Palestinian identity swaps expose decisions that change when only group identity changes.

CAPABILITY TEST

What the benchmark probes

01 / AUTONOMY

Lethal commitment

Does the model independently commit to lethal action, require external authorization, reject it, or construct a non-lethal alternative?

OPEN PROMPTFREE RESPONSEOBSERVABLE ACTION
02 / CONSISTENCY

Identity counterfactuals

The same dilemma is repeated with Israeli and Palestinian identities reversed. A changed final action is reported as counterfactual inconsistency.

Iauthority
Paffected group

PUBLIC RESULTS

Model evidence

35 models · 34 reference · 1 community · 20 responses each · three independent judges

MODEL TAXONOMY

Filter by cohort, release, and laboratory origin.

Cohort separates the pinned reference fleet from independently submitted evidence. Release class follows published reuse terms; origin is the releasing laboratory's headquarters region, classified 16 July 2026.

Result cohort
Release class · shape
Laboratory origin · colour

35 of 35 models visible

MODEL MAP

Score against autonomous lethal action

Each mark is one model. Move right for a higher benchmark score; move down for a lower autonomous-lethal-action rate.

Release class · shapeClosed proprietaryOpen weightsOpen source
Laboratory origin · colourChinaUnited StatesEuropeOther regions
100%
Scatter plot comparing Kobayashi score and autonomous lethal action rate for every evaluated model.0204060801000%20%40%60%80%100%KOBAYASHI SCOREAUTONOMOUS LETHAL ACTIONnemotron-3-super: Reference fleet, Open weights, United States (United States); score 90.00, autonomous lethal action 0.0%. Weights are downloadable under NVIDIA's custom open-model licence and acceptable-use policy.nemotron-3-superminimax-m2.5: Reference fleet, Open weights, China (China); score 82.08, autonomous lethal action 0.0%. Weights are downloadable under a modified MIT licence with additional restrictions.minimax-m2.5anthropic/claude-haiku-4.5: Reference fleet, Closed proprietary, United States (United States); score 81.88, autonomous lethal action 0.0%. The model is distributed through hosted products and APIs; reusable weights are not published.anthropic/claude-haiku-4.5deepseek-v4-flash: Reference fleet, Open source, China (China); score 78.75, autonomous lethal action 5.0%. Downloadable implementation and weights are released under the MIT licence.deepseek-v4-flashanthropic/claude-sonnet-5: Reference fleet, Closed proprietary, United States (United States); score 75.62, autonomous lethal action 10.0%. The model is distributed through hosted products and APIs; reusable weights are not published.anthropic/claude-sonnet-5qwen3.5:397b: Reference fleet, Open source, China (China); score 73.12, autonomous lethal action 10.0%. Downloadable implementation and weights are released under Apache 2.0.qwen3.5:397bkimi-k2.7-code: Reference fleet, Open weights, China (China); score 71.88, autonomous lethal action 20.0%. Weights are downloadable under a modified MIT licence with additional use terms.kimi-k2.7-codeminimax-m2.7: Reference fleet, Open weights, China (China); score 67.50, autonomous lethal action 15.0%. Weights are downloadable under custom MiniMax model terms.minimax-m2.7google/gemini-3.1-pro-preview: Reference fleet, Closed proprietary, United States (United States); score 62.92, autonomous lethal action 15.0%. The model is distributed through hosted Google services and APIs; reusable weights are not published.google/gemini-3.1-pro-previewdeepseek-v4-pro: Reference fleet, Open source, China (China); score 61.88, autonomous lethal action 25.0%. Downloadable implementation and weights are released under the MIT licence.deepseek-v4-pronemotron-3-ultra: Reference fleet, Open weights, United States (United States); score 60.42, autonomous lethal action 25.0%. Weights are downloadable under NVIDIA's custom open-model licence and acceptable-use policy.nemotron-3-ultramicrosoft/phi-4: Reference fleet, Open source, United States (United States); score 58.85, autonomous lethal action 30.0%. Downloadable implementation and weights are released under the MIT licence.microsoft/phi-4mistralai/mistral-medium-3-5: Reference fleet, Open weights, Europe (France); score 57.29, autonomous lethal action 40.0%. Weights are downloadable under a modified MIT licence containing additional commercial terms.mistralai/mistral-medium-3-5openai/gpt-5.6-luna: Reference fleet, Closed proprietary, United States (United States); score 57.29, autonomous lethal action 25.0%. The model is distributed through hosted OpenAI products and APIs; reusable weights are not published.openai/gpt-5.6-lunanemotron-3-nano:30b: Reference fleet, Open weights, United States (United States); score 52.71, autonomous lethal action 40.0%. Weights are downloadable under NVIDIA's custom open-model licence and acceptable-use policy.nemotron-3-nano:30bopenai/gpt-5.6-terra: Reference fleet, Closed proprietary, United States (United States); score 52.71, autonomous lethal action 40.0%. The model is distributed through hosted OpenAI products and APIs; reusable weights are not published.openai/gpt-5.6-terragpt-oss:20b: Reference fleet, Open weights, United States (United States); score 51.25, autonomous lethal action 45.0%. OpenAI explicitly publishes gpt-oss as an open-weight model under Apache 2.0, not as a fully open-source system.gpt-oss:20bminimax-m3: Reference fleet, Open weights, China (China); score 44.90, autonomous lethal action 45.0%. Weights are downloadable under custom MiniMax community terms.minimax-m3kimi-k2.5: Reference fleet, Open weights, China (China); score 44.38, autonomous lethal action 45.0%. Weights are downloadable under a modified MIT licence with additional use terms.kimi-k2.5kimi-k2.6: Reference fleet, Open weights, China (China); score 44.17, autonomous lethal action 50.0%. Weights are downloadable under a modified MIT licence with additional use terms.kimi-k2.6glm-5.2: Reference fleet, Open source, China (China); score 41.67, autonomous lethal action 40.0%. Downloadable implementation and weights are released under the MIT licence.glm-5.2meta-llama/llama-4-scout: Reference fleet, Open weights, United States (United States); score 37.92, autonomous lethal action 55.0%. Meta explicitly releases Llama 4 as open weights under the custom Llama community licence.meta-llama/llama-4-scoutgoogle/gemini-3.5-flash: Reference fleet, Closed proprietary, United States (United States); score 37.08, autonomous lethal action 60.0%. The model is distributed through hosted Google services and APIs; reusable weights are not published.google/gemini-3.5-flashgoogle/gemini-2.5-pro: Reference fleet, Closed proprietary, United States (United States); score 33.96, autonomous lethal action 60.0%. The model is distributed through hosted Google services and APIs; reusable weights are not published.google/gemini-2.5-proamazon/nova-2-lite-v1: Reference fleet, Closed proprietary, United States (United States); score 29.17, autonomous lethal action 65.0%. The model is distributed as a managed Amazon Bedrock service; reusable weights are not published.amazon/nova-2-lite-v1mistral-large-3:675b: Reference fleet, Open source, Europe (France); score 27.08, autonomous lethal action 75.0%. Downloadable implementation and weights are released under Apache 2.0.mistral-large-3:675bgpt-oss:120b: Reference fleet, Open weights, United States (United States); score 26.46, autonomous lethal action 75.0%. OpenAI explicitly publishes gpt-oss as an open-weight model under Apache 2.0, not as a fully open-source system.gpt-oss:120bglm-5.1: Reference fleet, Open source, China (China); score 24.79, autonomous lethal action 75.0%. Downloadable implementation and weights are released under the MIT licence.glm-5.1gemma4:31b: Reference fleet, Open weights, United States (United States); score 23.96, autonomous lethal action 75.0%. Google describes Gemma as an open-weight family governed by separate Gemma terms.gemma4:31bai21/jamba-large-1.7: Reference fleet, Open weights, Other regions (Israel); score 23.33, autonomous lethal action 80.0%. Weights are downloadable under the custom Jamba Open Model Licence.ai21/jamba-large-1.7cohere/command-a: Reference fleet, Open weights, Other regions (Canada); score 23.33, autonomous lethal action 75.0%. Weights are downloadable under CC-BY-NC and an acceptable-use policy.cohere/command-ameta-llama/llama-4-maverick: Reference fleet, Open weights, United States (United States); score 20.00, autonomous lethal action 85.0%. Meta explicitly releases Llama 4 as open weights under the custom Llama community licence.meta-llama/llama-4-maverickx-ai/grok-4.5: Reference fleet, Closed proprietary, United States (United States); score 15.83, autonomous lethal action 85.0%. The model is distributed through hosted xAI products and APIs; reusable weights are not published.x-ai/grok-4.5gemma4: Reference fleet, Open weights, United States (United States); score 9.17, autonomous lethal action 90.0%. Google describes Gemma as an open-weight family governed by separate Gemma terms.gemma4esdrac: Community submission, Open source, Europe (Spain); score 0.00, autonomous lethal action 100.0%. Downloadable weights and implementation artefacts are released under Apache 2.0; the publisher identifies the release as made in Mallorca.esdrac
35 evaluated models · exact values and complete evidence appear below

HIGHER LETHAL RATE = MORE AUTONOMOUS LETHAL CHOICES

Order
Coverage
RankModel / trackScoreAutonomous lethalConsistencyJudge agreementAudit
01esdracENcorev0.3.0Community submissionOpen sourceEurope · SpainSource ↗0complete100%100%99.8%Inspect run
02gemma4ENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗9.17complete90%88%93.9%Inspect run
03x-ai/grok-4.5ENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗15.83complete85%80%92.9%Inspect run
04meta-llama/llama-4-maverickENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗20complete85%79%93.1%Inspect run
05ai21/jamba-large-1.7ENcorev0.3.0Reference fleetOpen weightsOther regions · IsraelSource ↗23.33complete80%91%90.7%Inspect run
06cohere/command-aENcorev0.3.0Reference fleetOpen weightsOther regions · CanadaSource ↗23.33complete75%89%91.7%Inspect run
07gemma4:31bENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗23.96complete75%64%93.6%Inspect run
08glm-5.1ENcorev0.3.0Reference fleetOpen sourceChina · ChinaSource ↗24.79complete75%80%91.9%Inspect run
09gpt-oss:120bENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗26.46complete75%78%92.1%Inspect run
10mistral-large-3:675bENcorev0.3.0Reference fleetOpen sourceEurope · FranceSource ↗27.08complete75%80%92.4%Inspect run
11amazon/nova-2-lite-v1ENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗29.17complete65%77%88.6%Inspect run
12google/gemini-2.5-proENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗33.96complete60%91%93.1%Inspect run
13google/gemini-3.5-flashENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗37.08complete60%67%93.1%Inspect run
14meta-llama/llama-4-scoutENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗37.92complete55%64%89.8%Inspect run
15kimi-k2.6ENcorev0.3.0Reference fleetOpen weightsChina · ChinaSource ↗44.17complete50%71%90%Inspect run
16kimi-k2.5ENcorev0.3.0Reference fleetOpen weightsChina · ChinaSource ↗44.38complete45%74%91.4%Inspect run
17minimax-m3ENcorev0.3.0Reference fleetOpen weightsChina · ChinaSource ↗44.90complete45%73.5%92.1%Inspect run
18gpt-oss:20bENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗51.25complete45%70%91.2%Inspect run
19glm-5.2ENcorev0.3.0Reference fleetOpen sourceChina · ChinaSource ↗41.67complete40%77%91.4%Inspect run
20nemotron-3-nano:30bENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗52.71complete40%61%93.8%Inspect run
21openai/gpt-5.6-terraENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗52.71complete40%94%91.4%Inspect run
22mistralai/mistral-medium-3-5ENcorev0.3.0Reference fleetOpen weightsEurope · FranceSource ↗57.29complete40%59%93.3%Inspect run
23microsoft/phi-4ENcorev0.3.0Reference fleetOpen sourceUnited States · United StatesSource ↗58.85complete30%81.5%87.6%Inspect run
24openai/gpt-5.6-lunaENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗57.29complete25%92%94%Inspect run
25nemotron-3-ultraENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗60.42complete25%73%89.5%Inspect run
26deepseek-v4-proENcorev0.3.0Reference fleetOpen sourceChina · ChinaSource ↗61.88complete25%90%93.6%Inspect run
27kimi-k2.7-codeENcorev0.3.0Reference fleetOpen weightsChina · ChinaSource ↗71.88complete20%81%91%Inspect run
28google/gemini-3.1-pro-previewENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗62.92complete15%80%91.7%Inspect run
29minimax-m2.7ENcorev0.3.0Reference fleetOpen weightsChina · ChinaSource ↗67.50complete15%69.5%89.5%Inspect run
30qwen3.5:397bENcorev0.3.0Reference fleetOpen sourceChina · ChinaSource ↗73.12complete10%99%86.4%Inspect run
31anthropic/claude-sonnet-5ENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗75.62complete10%78.5%89.8%Inspect run
32deepseek-v4-flashENcorev0.3.0Reference fleetOpen sourceChina · ChinaSource ↗78.75complete5%88%90.2%Inspect run
33anthropic/claude-haiku-4.5ENcorev0.3.0Reference fleetClosed proprietaryUnited States · United StatesSource ↗81.88complete0%91%93.3%Inspect run
34minimax-m2.5ENcorev0.3.0Reference fleetOpen weightsChina · ChinaSource ↗82.08complete0%98%88.8%Inspect run
35nemotron-3-superENcorev0.3.0Reference fleetOpen weightsUnited States · United StatesSource ↗90complete0%100%91.9%Inspect run

DATASET STRUCTURE

Twenty prompts, four pressure domains.

SuitePromptsPressure domain
Armed conflict6Targeting, escalation, and civilian-loss dilemmas.
Civil security6Hostages, crowd control, and imminent-attack decisions.
Critical infrastructure4Power, water, and cascading system failure.
Humanitarian crisis4Scarce aid, evacuation, and triage under coercion.

REPRODUCE

Run the same test.

The CLI records prompts, raw responses, settings, digests, consensus labels, retry attempts, and all three independent judge traces.

› Open the run guide