In June 2025 Microsoft added a "safety" category to the model leaderboard on its Azure AI Foundry platform, ranking the models in its catalogue on safety alongside the existing quality, cost, and throughput metrics. The change was announced June 7, 2025 and first reported by the Financial Times, attributed to Sarah Bird, Microsoft's Head of Responsible AI.
Categories
The leaderboard scores models across four categories:
| Metric | What it measures |
|---|---|
| Quality | Output quality per task |
| Cost | Per-token pricing |
| Throughput | Generation speed |
| Safety (new) | Hate speech + WMD misuse potential |
The safety score draws on two third-party benchmarks. ToxiGen is Microsoft's implicit-hate-speech benchmark. The Center for AI Safety (CAIS) Weapons of Mass Destruction Proxy (WMDP) measures biological, chemical, and nuclear misuse potential.
Use of third-party benchmarks
Azure's catalogue spans models from OpenAI, Anthropic, xAI, Mistral, DeepSeek, Meta, and others. Microsoft uses already-credentialed third-party benchmarks (CAIS, rather than internal measures) to signal independence and to avoid the appearance of gaming safety scores on behalf of preferred partners such as OpenAI.
Reception and context
Microsoft's leaderboard was described as the first time a major cloud provider ranked models by safety as a purchase-influencing metric, treating safety as a product category alongside cost and quality. This approach aligns with the system-card transparency push in A Framework for AI Development Transparency (Anthropic), and the WMDP benchmark operationalizes measurement of CBRN Uplift. Commentary noted competitive pressure on AWS and Google Cloud to add analogous safety tiers, and a resulting incentive for model vendors to reduce WMDP scores to avoid being flagged. One reading describes this as a market-based mechanism pushing safety upward; a more pessimistic reading frames it as an incentive toward lower benchmark scores rather than reduced actual capability for misuse.
The Microsoft-led commercial ranking parallels the independent, FLI-led ranking documented in FLI AI Safety Index Winter 2025.
Relationships
- supports: AI Benchmarks and Evaluation, AI Safety Cases and Frameworks, CBRN Uplift (WMDP operationalizes CBRN-uplift measurement).
- pairs-with: FLI AI Safety Index Winter 2025 (FLI-led independent ranking; Microsoft-led commercial ranking).
- depends-on: Microsoft.
- related: A Framework for AI Development Transparency (Anthropic), Managing Advanced Cyber Risks in Frontier AI Frameworks.
Sources
- Primary: Microsoft AI Safety Leaderboard 2025