Cloudflare AI Gateway binds requests to verified user identities
For the first time, organizations can answer who sent what to which AI model without reconstructing shared API key usage.
TL;DR
Cloudflare opened beta for identity-aware AI Gateway, which binds every AI request to a verified user identity through Cloudflare Access. Organizations using Okta, Entra, or other SAML providers can now attribute each model call to an individual, ending the shared-API-key blind spot that made AI governance aspirational. User Insights, now generally available at no added cost, builds behavioral baselines per user and flags anomalies in spend, volume, or tool usage. Per-user spend limits are included.
Most organizations running AI workloads today can't answer a basic governance question: who sent that prompt? Shared API keys, the default for most enterprise AI setups, collapse all usage into one undifferentiated stream. A spike in spend or a prompt containing proprietary code has no name attached to it. Cloudflare's identity-aware AI Gateway, now in open beta, addresses this at the infrastructure layer: every request routed through the gateway carries a verified user identity from the organization's SAML provider.
The mechanism is straightforward. Cloudflare Access sits in front of AI Gateway, authenticates the user, and passes a verified cf.user_id into request metadata. From there, logs, analytics, and spend become attributable to individuals, not to a shared key. Per-user spend limits let organizations cap consumption by person or group. The companion feature, User Insights, builds behavioral baselines per user and flags anomalies: a 10x usage spike, a new model being accessed, an agent going off-pattern. Flexport's staff security engineer confirmed the real-world pain point in Cloudflare's announcement: shared API keys made it "almost impossible to tell who is using an AI service."
For compliance teams evaluating this, the limitation is scope. Identity-aware AI Gateway is an observability and budgeting layer, not a DLP or content-filtering solution. It tells you who called which model and how much they spent; it doesn't inspect prompts for CUI or block sensitive data from leaving the perimeter. Organizations that need prompt-level DLP will still need complementary controls, Cloudflare sells those separately through its CASB and DLP products. And the identity-binding feature is in open beta, so production deployments should factor for the usual beta caveats. But for the narrow, high-value problem of attribution (knowing who did what with AI) this is a genuine step past the shared-key blind spot that has made AI governance aspirational rather than operational.
Published ·Deep Fathom