Scale appoints Francis deSouza as the new CEOLearn more

To Steer the AI Frontier, Washington Must Build Its Testing Power

By Francis deSouza·September 25, 2026·4 min read
White Scale logo against a blue background with gradient bars

Much of today's debate about AI risk rests on speculation. Some of the loudest warnings have little grounding in science, yet they drive policy as if they were proven.

Getting this wrong carries costs in both directions. Moving recklessly would be a mistake, but slowing down out of fear would be just as wrong. Every week lost is a week of slower progress toward new cancer treatments, stronger American national security, and scientific breakthroughs.

The way forward is evidence.

AI is a central topic at this week's UN General Assembly as world leaders discuss how to coordinate their efforts internationally. Any agreement they reach will only be as strong as the evidence beneath it. That means governments need to get specific about testing: what are the benchmarks, who verifies the results, and what happens when serious risks are found.

The need for a stronger approach to testing and evaluation is especially urgent in Washington. The federal government’s ability to evaluate frontier models has not kept pace with the companies building them. That gap needs to close. Policymakers need independent testing to reliably separate demonstrated risks from imagined ones. It’s like asking a doctor to prescribe treatment without access to the right diagnostic tests. You cannot regulate what you cannot measure.

Today, most frontier model testing happens inside the labs themselves. These companies have a moral responsibility and a strong commercial incentive to test their systems and release safe products. But on questions of national security and public safety, Washington cannot let the labs grade their own homework. The government needs its own evidence on cyber, biological, autonomy-related, and other emerging risks.

Building an effective government testing system requires three things:

  • A clear mandate: Designate an accountable official to coordinate testing across federal agencies, set government-wide evaluation priorities, mobilize evaluation expertise across the public and private sectors, and ensure serious risks are identified and mitigated.
  • Pre-deployment evaluation: Strengthen the government’s ability to evaluate frontier models before deployment by working with model developers to secure access for testing and reassess national security risks as capabilities evolve and new uses emerge.
  • Dedicated capacity: Commit sustained government funding to independent public-private testing, including the testing tools, resources, and technical expertise needed across national security domains.

Putting those priorities into practice will take continued investment. The President requested $27 million for CAISI, the government's frontier AI testing center, in fiscal year 2027, while House appropriators proposed up to $15 million. Given the magnitude and speed of AI progress, these funding levels are like managing JFK air traffic control from a small-town control tower.

While the government needs more resources, it does not have to build everything from scratch. At Scale, we’ve spent years developing the expertise and building the infrastructure to test advanced models on the hardest problems. We have partnered with government AI evaluation bodies in the United States, the United Kingdom, Korea, and Singapore to develop benchmarks and expand testing capacity, and we are accelerating that work, with more to share in the coming weeks. In addition, we continue to assist national security partners in building cyber evaluations that measure what models can do against real-world threats.

Our own cyber research has also found security weaknesses that standard tests might miss. Take AI agents that handle tasks like email and banking on a person’s behalf. Our researchers found that when an agent pauses during one of these tasks to ask for clarification, an attacker can slip harmful instructions into the answer. Most models tested were vulnerable to this form of attack. For several models, attacks succeeded about 10 times more often when malicious instructions appeared in answers to their questions than when hidden in emails, documents, or webpages they read. These attacks can also go unnoticed because an agent can complete the user’s task correctly while carrying out the attacker’s instructions at the same time.

We’ve also found the opposite problem. AI can refuse to help people defend their systems. In a national cyber defense competition, where students defended live computer systems against professional attackers, AI models refused to help the defenders 12 percent of the time, and more than 40 percent of the time when they asked how to lock down their systems. Without testing built for these scenarios, policymakers cannot know what America's banks, power grids, and hospitals need to be ready for.

That is the kind of evidence Washington needs. With the capacity to direct outside experts and act on what they find, it can address real risks before they become crises and clear the way for everything AI promises.

The time to act is now. The nation that can measure the frontier will be the one that steers it.

Ready to break through your data bottleneck?

Scale's team will match your project to the right experts, fast.