ai-compliancetrade-pressNewsThe Broadside2 min read

CSIS pitches AI benchmarking as EO deliverables stay dark

The June 2 executive order required a voluntary framework by August 1 and a classified benchmarking process, neither has surfaced publicly, and a FOIA suit is now the only release clock.


TL;DR

CSIS is pushing federal investment in AI benchmarking (including university-hosted policy labs and public-private partnerships) as a way to build trust in frontier models before they reach national-security workflows. The timing isn't accidental. President Trump's June 2 EO on frontier AI required a voluntary framework for pre-release model access by August 1; that framework hasn't been made public. A separate classified benchmarking process, led by Treasury with NSA determining covered models, also hasn't surfaced. The nonprofit Protect Democracy Project sued ONCD, Commerce, Treasury, and the White House Office of Science and Technology on September 1 under FOIA; the government agreed September 15 to release more documents by October 30.

The Center for Strategic and International Studies made its case this week: AI benchmarking shouldn't live exclusively inside the classified channels the Trump administration has been building since June. In a September 17 post, CSIS fellows Benjamin Jensen and Yasir Atalan argued that independent institutions (universities, nonprofits, federally funded policy labs) should set evaluation standards that establish "the foundation of trust in AI systems," including for military use cases.

They're making this argument into a vacuum. EO 14409, signed June 2, tasked Treasury, NSA, and CISA with developing a classified benchmarking process to assess frontier models with advanced cyber capabilities and to set a threshold that triggers government pre-release testing access. That process hasn't been described publicly. The EO also required a voluntary framework (due August 1) to determine which models get covered, govern access for testing, and encourage collaboration with "trusted partners." That framework hasn't appeared either.

On September 1, the nonprofit Protect Democracy Project sued four agencies under FOIA to force release of the pre-release framework and information on participating companies. The government agreed on September 15 to release additional documents, with fewer redactions, by October 30. That's the only deadline anyone can count on right now.

Jensen and Atalan's proposals (NSF-administered research grants, a token-use levy to fund public-interest research, NIST's CAISI and NSF's X-Labs as partnership models) amount to an alternative theory of how benchmarking should be done: distributed, transparent, and academic, not classified and interagency. Whether that theory has any purchase depends on what the administration's framework actually says when it surfaces. Four months after the EO, no one outside the government knows.

The CSIS post landed the same week Anthropic CEO Dario Amodei called for slowing frontier AI development to allow safety collaboration, a position OpenAI CEO Sam Altman immediately endorsed. Jensen and Atalan framed the moment as one where "governments have struggled to keep up with the private sector in terms of regulating technology and commerce." Their remedy is federal money routed through universities and nonprofits, institutions whose budgets, they noted, the NSF has seen cut by more than half in recent years.


Published ·Deep Fathom