Back to Blog·AI / LLM

No AI Lab Scored Above a C+ on Safety in 2026 — What That Means If You're Choosing a Vendor

The Future of Life Institute's Summer 2026 AI Safety Index graded nine major AI labs. Anthropic came out on top — at a C+. Here's what the grades actually measure, and how much weight they should get in a vendor decision.

Majid Hussain· Founder & CEO, DIGIT7 min read

The Future of Life Institute's Summer 2026 AI Safety Index graded nine leading AI companies across 37 indicators in six domains. The best grade any company earned was a C+. That's worth sitting with before we get to the individual scores.

The Scores

Anthropic topped the index at C+ (2.66), OpenAI and Google DeepMind followed at C (2.28 and 2.01 respectively), Meta scored D+, and xAI, DeepSeek, and Mistral all received failing F grades — with Z.ai and Alibaba Cloud at D-. The six domains assessed: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information sharing. Anthropic led five of the six domains; OpenAI led on risk assessment specifically.

The Finding That Should Concern Enterprise Buyers More Than the Letter Grades

Reviewers flagged that Anthropic, OpenAI, Google DeepMind, and Meta have all weakened earlier public pledges to pause development at certain danger thresholds — reviewers characterized this plainly as "moving the goalposts." Separately, all four of those companies, which had previously banned military applications of their models, have gradually reversed that position and now pursue defense partnerships, joining xAI and Mistral in that space. A grade is a snapshot; a trend of walking back safety commitments is a pattern, and patterns predict future behavior better than a single index score does.

No Company Scored Above a C- on Existential Safety

Across all nine companies, no one scored better than C- on the existential-safety domain specifically, and most scored D or below — Anthropic's D+ was the single best grade in that category. This is the domain most relevant to frontier-model risk specifically (loss of control, catastrophic misuse) rather than near-term product safety, and it's the one every company in the industry is doing worst on.

How Much Weight Should This Actually Get in a Vendor Decision?

A meaningful amount, but not as a sole deciding factor. The Safety Index measures organizational governance and existential risk posture — genuinely important for understanding a vendor's trajectory and public commitments, but it doesn't tell you whether a specific model will hallucinate on your specific use case, or whether the vendor's guardrails and evaluation tooling are strong enough for your application. We cover what to actually ask an AI vendor beyond public safety scores — data handling, evaluation methodology, escalation paths — in a separate guide, because those questions matter just as much for your specific project.

What We'd Actually Tell a Client

Use the Safety Index as one input on vendor trajectory and public commitment, not a pass/fail gate — a C+ leader and a D+ competitor can both be viable depending on your actual use case and how much of the risk surface your own architecture (guardrails, evaluation, human escalation) is responsible for covering regardless of which model sits underneath it. The index measures the company; your own implementation still determines most of the risk your users actually experience.

If you're weighing AI vendor choice against safety, governance, or public trust considerations, reach out at info@digit.com.pk — we'll help you separate what the vendor's own safety posture covers from what your implementation needs to cover regardless of vendor.

#AIsafetyindex2026#AnthropicOpenAIsafetycomparison#enterpriseAIvendorsafety#FutureofLifeInstituteAIgrades#digitpk#digit#digitio
Share

Related Articles

Built by DIGIT

Need help building something like this?

DIGIT has shipped 1,000+ projects across web, mobile, AI and cloud. Let's talk about yours.