methodology
Every number links to how it was produced.
The full methodology publishes with the first benchmarks: arms and controls, grading, confidence intervals, Trust Score weights, and the integrity wall between money and scores.
- Lift = pass rate with skill minus without, 95% CI, paired bootstrap
- Same model, runtime and image across arms; fresh sandbox per trial
- Trust Score: security 40% · quality 20% · license 15% · maintenance 15% · community 10%
- Scoring changes are logged in a public changelog
Want in first?