AI Model Leaderboard for Nonprofits

A ranking of frontier AI models across five pillars weighted for nonprofit priorities: capability, alignment & honesty, cost & efficiency, bias & harm resistance, and security posture.

This leaderboard is an aggregate of independent public benchmarks from reliable sources. It's not based on tests performed by our organization.

The default overall score is calculated with these weights: capability 40%, alignment and honesty 15%, cost and efficiency 15%, bias and harm resistance 15%, and security 15%.

If your organization has different priorities (for example, giving more weight to cost, alignment, bias or security), use the "Change priorities" buttons to re-rank the models, or compare the pillar columns directly.

Expensive models usually require more energy and resources, so higher cost usually translates also as higher environmental impact.

Click a model name to see score details. Click a column header to sort.

Change priorities:
# Model Provider License Overall score Capability Cost Alignment Bias & harm Security
Greener means higher score, redder means lower score. (!) score proxied from the previous model version. (≈) pillar imputed at the set median.
Methodology, sources & limitations

How the overall score works

Each raw metric is normalized to a 0-100 scale with min-max within this tracked set of models (cost per task inverted, so cheaper = higher). Pillar scores are the mean of their normalized metrics. The overall score is a weighted mean of the five pillar scores:

The weights are a documented hypothesis derived from nonprofit sector priorities (cost pressure, donor/constituent trust, bias concerns) and are open to challenge. Because every pillar score is published separately, any organization can re-rank the table with its own weights using the open data file. The "Change priorities" buttons above recompute the overall score immediately in your browser (the selected pillar at 40% and the rest at 15%; capability is selected initially, matching the default weights).

Proxied scores (⚠)

When a newly released model is not yet evaluated by a source, its score is inherited from the nearest previous model from the same provider (for example GPT-6.1 Sol inherits GPT-6 Sol's Enkrypt ratings). Every proxied score is marked with ⚠ and a callout naming the source model. Proxied models can outrank fully measured ones.

Imputed pillars (≈)

When no data and no predecessor exist for an entire pillar (e.g. HY4 Preview is not benchmarked by Artificial Analysis at all), that pillar is imputed at the set median so the composite stays comparable. A missing pillar is neither rewarded nor punished. Imputed pillars are marked ≈ and listed per model; models with imputed pillars carry the partial badge, so read their rank with caution.

Known gaps, stated plainly

Sources

Not an endorsement

This index summarizes third-party measurements for decision support. It is not a safety certification, not legal advice, and not an endorsement of any vendor.