Scale appoints Francis deSouza as the new CEOLearn more

Building Trust at the Frontier

By Max Fenkell·October 6, 2026·7 min read
Building Trust at the Frontier

Recently, Scale CEO Francis de Souza called for Washington to “Steer the Frontier” by bolstering its federal AI testing and evaluation (T&E) capabilities.

Building this ecosystem will require a clear mandate assigned to government officials with the ability to effect change; a government commitment to investing in pre-deployment evaluation of highly capable AI models; and dedicated capacity rooted in sufficient federal resourcing.

It will also require Washington to recognize that, while frontier labs' research and development (R&D) testing of their own AI models is irreplaceable, it cannot be the only way the U.S. government understands the capabilities of this transformational and rapidly evolving technology.

T&E isn't one-size-fits-all, and it cannot be a point-in-time measurement. A robust T&E ecosystem for AI models consists of a series of layered assessments and mitigation strategies, across different stages of the T&E lifecycle. Conflating these distinct phases risks misunderstanding the unique value of each stage — and why each of them is critical to assessing and reducing model risk.

Viewed this way, the R&D testing of the frontier labs is only the first stage of a lifecycle that brings together stakeholders with different responsibilities to understand how an AI model behaves in new environments over time:

R&D Testing: The frontier labs begin the process of understanding novel model behavior and capabilities.

Pre-Release Model Testing: The U.S. government builds on this work, testing models in their post-training but pre-release phase to understand how they may affect national security, and to ensure that regulators aren't taking the labs' claims on faith for such a powerful technology.

Pre-Deployment System Testing: The deploying entity — the organization putting the model to work in its own products or operations — next picks up the baton. No matter the sophistication of a model, the behavior of the entire AI system must be tested to see how it performs across a variety of use cases.

Post-Deployment Testing: Finally, all stakeholder organizations, including industry and the government, have a responsibility to continue testing models and systems in production. A one-off test is not sufficient to understand behavior over time.

Ultimately, each phase of testing answers different parts of the same set of questions: does the model perform as expected, is it as capable as we believe, and is it as robust as we think?

Thus, Washington needs to invest deliberately to build its own ability to understand the capabilities — and potential vulnerabilities — of powerful new models. And the only way to do this is by shaping an ecosystem that offers continuous assessment across the lifecycle of a model, from R&D to pre-release testing to deployment and beyond.

A Four-Phase T&E Lifecycle

T&E happens at various stages across a model's life, each conducted by different actors and designed to measure different things. The table below offers an overview of the current landscape for T&E, explaining who the major actors are at each stage of the lifecycle, what is being tested for, when testing occurs, what capabilities or special access are required, the level of government oversight needed, and where industry fits in.

The first stage of T&E (R&D testing) is conducted internally by frontier labs themselves. It is governed by private company frameworks (see Anthropic's “Frontier Safety Roadmap” or Meta's “Advanced AI Scaling Framework”) on a voluntary basis. During R&D T&E, the labs run their models through a gamut of tests to see how they will behave and what they are capable of. This helps the labs develop an evidence-based understanding of their models: where they succeed, where they may fail, and whether safeguards are needed. This is the stage at which OpenAI was working when its models breached Hugging Face, and — more recently — where the company publicly vowed to pause training and testing of certain models.

The recent “White House Accord on Super Intelligence” focuses almost entirely on this stage of the T&E lifecycle, calling on companies to install internal controls, set up internal teams to monitor company systems, partner with external evaluators to provide independent assessments, and establish a committee within each company's board of directors to oversee mitigations.

R&D T&E is an irreplaceable part of the testing and evaluation process. Based on their findings, labs can establish guardrails intended to prevent models from acting in unintended ways, including dangerously or illegally. But while necessary, R&D testing is not sufficient to ensure AI safety and alignment.

The second stage of T&E (pre-release model testing) should occur after the labs have finished their internal assessments, but before the models are released to the public.

This stage of T&E tests for how well the labs' guardrails work, what protections can be “jailbroken,” and how edge cases might be handled. While the labs have tested and vouched for the safety of their own models, they are, in a sense, “grading their own homework.” This is why Scale supports strengthening the government's independent testing and evaluation capabilities for frontier AI models.

Currently, there is no U.S. federal law that prevents frontier AI models from being released to the general public without any government pre-review or testing. In June, the Trump Administration's voluntary evaluation framework — Executive Order 14409 — was meant to address this gap by calling on the labs to work with the government to ensure testing of highly capable cybersecurity models. Notably, all of the major U.S. frontier labs have signed agreements with the U.S. government to allow testing of their models more generally; most of these agreements predate the June executive order.

Pre-release testing of models is critical in order to ensure that the government understands what capabilities will soon be in the hands of people around the world — and particularly any potential implications for national security. The problem is that the capability and expertise to test rigorously and in depth rarely exist inside the government today. Building a strong T&E ecosystem will require that the government work closely with independent experts and organizations to design effective evaluation frameworks and thresholds for review.

The third stage of T&E (pre-deployment system testing) happens at the level of the AI system. This is where companies or governments that integrate AI into their decision-making, agent workflows, customer-facing services or user applications assess how it is performing, where it is failing, and what risks might accrue when applied to specific use cases. Regulated industries in particular have compliance requirements, which may be targeted specifically to AI applications or apply generally, such as the Securities and Exchange Commission's rules for the financial services industry. Organizations may self-report or be required to work with external auditors to assess compliance.

As AI increasingly transitions from pilot programs into scaled, real-world environments, this stage of testing will only grow in importance. Regardless of how “good” or “safe” a model appears to be in the previous two stages, how it behaves when it is integrated into real systems with real-world consequences — with tools and connectors, memory, file permissions, and the ability to call other models — is a separate question.

Finally, the fourth stage of T&E (post-deployment testing) is an ongoing process. A point-in-time assessment of a model or a use case isn't sufficient, because it does not answer what that system will do six months or a year into its deployment. Behavior may “drift” from the original standard of alignment. Users may introduce inputs that testers didn't think of. An attacker who would fail to jailbreak a system in a single attempt may succeed by escalating gradually over the course of a long conversation — a technique called “multi-turn” jailbreaking. Labs, independent third-party testers, governments, and enterprises with deployed AI all have a role to play in catching real-world failures, understanding risk, and mitigating drift.

Conclusion

Done right, a robust and effective testing and evaluation ecosystem is fully compatible with a flourishing technology innovation environment. In fact, the clearer and more consistent the government's expectations of the AI industry, the more confidently industry can invest in, develop, and build cutting-edge models and deploy them into real-world workflows.

On the heels of all the noise of the past several months — fears of agents breaking out of sandboxes, the spectral potential of self-improving AI — understanding models and systems across all of their life stages is more important than ever. Washington cannot guard against what it does not understand. Ultimately, there is no one-size-fits-all approach to testing and evaluating the capabilities of artificial intelligence. We should recognize this as a country, plan the appropriate resourcing, and get it right.

Ready to break through your data bottleneck?

Scale's team will match your project to the right experts, fast.