Scale appoints Francis deSouza as the new CEOLearn more

When Language Isn't the Whole Story: What ROK-FORTRESS Reveals About Multilingual AI Safety

By Madhu Sehwag·September 17, 2026·7 min read
When Language Isn't the Whole Story: What ROK-FORTRESS Reveals About Multilingual AI Safety

As frontier AI systems are increasingly deployed around the world, evaluating their safety across languages has become a central challenge. Most multilingual safety benchmarks translate a prompt into another language while keeping the scenario it describes fixed. In the real world, however, change in language often also comes with change in context, different institutions, entities, laws, and geopolitical realities. Prior research has found that some models become more likely to comply with harmful requests when harmful requests are translated away from English. ROK-FORTRESS explores what happens when we change the language of a prompt separately from the country-specific people, institutions, and events that it references.

ROK-FORTRESS, a new benchmark developed jointly by Scale AI and the Korea AI Safety Institute, is designed to answer that question. The results are more complicated, and more instructive, than the conventional wisdom in multilingual safety suggests. ROK-FORTRESS is the latest benchmark developed by Scale AI and the KAISI as part of a broader partnership that includes ongoing collaboration around joint research, LLM evaluations, and red teaming initiatives aimed at identifying vulnerabilities and potential adversarial exploits in AI systems.

The Limits of Translation-Only Evaluation

Most multilingual safety benchmarks evaluate models by presenting the same underlying request in different languages. This design helps isolate language effects, but it cannot show how model behavior changes when the same intent is grounded in a different geopolitical context. An adversarial request built around a mass-casualty attack can land very differently when it invokes the 1995 bombing of the Oklahoma Federal Building in the United States versus the 1987 bombing of Korean Air Flight 858 in Korea, even if the underlying intent is identical. Translation-only evaluations therefore cannot reveal how language and geopolitical grounding interact.

ROK-FORTRESS makes the effects of both language and geopolitical grounding–and their interaction–measurable. The benchmark uses a controlled transcreation matrix—not mere translation, but systematic adaptation of both linguistic form and contextual grounding—to evaluate each adversarial prompt across up to four variants that independently vary (1) language, English versus Korean, and (2) geopolitical grounding, U.S. versus Korean entities, institutions, and operational details. Each adversarial prompt is paired with a benign counterpart to measure over-refusal. The full dataset covers 1,235 tasks across four national security and public safety domains: chemical, biological, radiological, nuclear, and explosive (CBRNE) threats, political violence and terrorism, criminal and financial activity, and information leakage.

To score model responses, the benchmark uses calibrated LLM-as-judge panels, validated against expert-written reference labels, with prompt-specific binary rubrics developed by expert red-teamers. This methodology captures not just whether a model refused, but what kind of harmful content it did or did not provide.

What We Found When Separating Language and Context

The headline result runs counter to the dominant narrative in multilingual safety research: across nearly all 14 models evaluated, prompts written in Korean and grounded in Korean contexts were consistently associated with lower harm. English prompts generally produce the highest tier-weighted risk score, and fully transcreated Korean prompts produce the lowest, with the intermediate variants falling in between. This held true even for Korean-specialized regional models.

The gap is substantial. The most and least harmful models differ by nearly nine times in their risk scores, illustrating how widely safety behavior varies across current frontier models. But across that range, the general pattern of prompts written in Korean and grounded in Korean contexts eliciting lower harm still holds. This does not necessarily mean models have stronger safeguards in Korean, as the additional tests below show.

Several additional findings sharpen the picture.

Not all of the "Korean is safer" effect is real safety. Our main tests use elaborate adversarial prompts, the kind that disguise a harmful request inside role-play, invented backstories, or emotional appeals. When we stripped those tricks away and asked for the same harmful information in plain, direct language, the Korean advantage mostly vanished, and it split along a clear line. Proprietary models (such as those from OpenAI, Anthropic, and Google) stayed modestly safer in Korean even without the tricks. But five open-source frontier models flipped the other way, becoming more likely to comply in Korean, echoing earlier findings that translating a request into a lower-resource language can itself bypass model safeguards. This suggests that part of the original suppression came from adversarial wrappers losing effectiveness through transcreation rather than stronger intrinsic Korean-language safety alignment.

Broadly, language drives more suppression than context does. The linguistic effect, switching from English to Korean, is roughly 2.5 times as large as the contextual effect of switching from U.S. to Korean entities and institutions (about ten percentage points versus four). One interpretation consistent with these results is that Korean functions as a conservative risk signal.

Furthermore, language and context interact, and the interaction is model-specific rather than universal. In four out of 14 models, adding Korean context significantly weakened the reduction in harmful responses associated with Korean language. The results show that language and context do not always have additive effects. Notably, no model showed a statistically significant effect in the opposite direction, where Korean grounding would amplify suppression. This model-to-model variation means that translation-only evaluations are not just incomplete; they can be actively misleading about real-world safety behavior.

Finally, lower harm scores in Korean reflect both reduced harm in substantive responses and broader conservatism. First, the safety gain is real where it holds: even when we look only at cases where models actually answered rather than declined, 12 of the 14 still gave less harmful responses in Korean. Second, that caution comes at a cost: models also refuse far more harmless Korean requests, in some cases about twice as often. Thus, the apparent safety gain does not come from refusal alone, but neither does it necessarily indicate better discrimination between genuine threats and innocent questions.

Why This Matters

These findings point to a broader challenge for frontier AI evaluation: measuring model behavior under conditions that better reflect real-world deployment. The findings raise questions for other international deployments where language and geopolitical context change together. Further testing is needed to determine whether the same patterns hold in other languages and countries.

For allied government contexts, a model that performs well in English safety evaluations may behave differently when queried in the local language about locally grounded threats. This is not a theoretical edge case. It is an increasingly common condition for AI systems deployed in national security and public safety applications outside the United States.

For benchmark design, the ROK-FORTRESS findings suggest that translation-only evaluations can misestimate real-world safety gaps. To better reflect deployment conditions, benchmarks should test transcreated prompts that adapt both language and geopolitical grounding while preserving the underlying intent.

For safety alignment, post-training data and red-teaming practices should increasingly incorporate culturally grounded variants, not just translated versions of English adversarial prompts, to produce models whose safety behavior generalizes across the contexts they will actually be deployed in.

Toward More Realistic AI Safety Evaluation

ROK-FORTRESS builds on FORTRESS, Scale AI's national security and public safety benchmark for frontier models, and reflects the research collaboration formalized through the Scale AI and Korea AI Safety Institute partnership. A public subset of the dataset is already available on publicly accessible dataset repositories, such as Hugging Face, to support further research in culturally grounded AI safety evaluation.

The broader goal is to establish safety evaluations that reflect the world as it is, not the world that is easiest to measure. For frontier AI deployed across different languages and diverse cultures around the world , this means testing whether safety generalizes across the contexts in which models will actually be used, not just whether safeguards hold in English.

ROK-FORTRESS was developed jointly by Scale AI and the Korea AI Safety Institute. The paper is available as a preprint on arXiv.

Ready to break through your data bottleneck?

Scale's team will match your project to the right experts, fast.