Methodology and evidence limits
Updated October 2, 2026 · Formula documented from the current scoring implementation.
Ranking order
The default leaderboard sorts by Pulse Score, highest first. Ties use newer release date, then model name, then slug. Scores are rounded to one decimal before ranking. Missing scores are placed after scored models. Filters and alternative sorts retain the global rank badges; row position under another sort is not a new global rank.
The static fallback is a snapshot sorted by its displayed scores. It can lag the interactive catalog. Equal scores in that snapshot retain their stored order when release dates are unavailable.
Text model formula
Each component is converted to a score out of ten. Multiply each available component by its base weight, add the products, then divide by the sum of included weights. Round to one decimal. Missing optional fields are omitted; their weights are not treated as zero-valued results.
| Component | Base weight | Calculation |
|---|---|---|
| Quality | 40% | First nonzero field: LLM Stats score, quality score, or Artificial Analysis intelligence index. LLM Stats uses 80 as the normalization ceiling; other fields use 100. Divide by the ceiling, multiply by 10, cap at 10. A missing quality signal leaves the text score unavailable. |
| Recency | 14% | Age bands in days: ≤14: 10; ≤45: 9.8; ≤90: 9.2; ≤180: 8.2; ≤365: 6.5; ≤730: 4; older: 1.5. Unknown dates receive 3. |
| Launch boost | 8% when positive | Age ≤30 days: 10; ≤60: 8.5; ≤120: 6; ≤180: 3.5. Otherwise omit. |
| Value | 12% when price exists | Use the mean input/output price per million tokens, or the available price. Let P=max(0,1−log(1+price)/log(26)); value=min(quality×(0.60+0.40×P),10). |
| Speed | 10% when positive | min(tokens per second / 200 × 10, 10). |
| Latency | 6% when positive | (1−min(time to first token in seconds / 15,1)) × 10. |
| Reasoning / coding | 5% each when positive | min(reported score / 100 × 10,10). |
Recency uses the current calculation time. Quarter-only dates use the fifteenth day of the first month of that quarter; year-only dates use July 1. Future dates currently receive the newest age band. These are scoring conventions, not benchmark evidence. A score can change as time passes even without new test results.
Other categories
Other categories use category-specific metrics, plus recency at a base weight of 15% and a positive launch boost at 8%. Available weights are normalized in the same way. Their scores should not be compared directly with text model scores.
| Category | Base metrics and weights |
|---|---|
| Images | ELO 40%, generation time 25%, resolution 15%, cost per 100 images 20% |
| Video | Fluidity 40%, duration 20%, frame rate 20%, cost per minute 20% |
| Embeddings | MTEB 45%, dimensions 15%, token limit 15%, cost per million tokens 25% |
| Rerankers | Hit rate 40%, latency 30%, context 15%, cost per thousand requests 15% |
| Speech generation | Voice realism 40%, latency 25%, language coverage 15%, cost per million characters 20% |
| Speech recognition | Word error rate 50%, cost per thousand minutes 30%, language coverage 20% |
| 3D | Topology quality 40%, generation time 25%, export formats 15%, cost per model 20% |
Subjective labels such as voice realism or topology quality require a published rubric to be reproducible. A number without that rubric is not an independently verified measurement.
Evidence required per result
- A direct source URL supporting the exact metric, plus provider or independent evaluator attribution.
- The exact model version or snapshot, benchmark version, task set, and run date.
- Prompting method, reasoning budget, tool access, sampling settings, number of trials, and scoring rules.
- For latency and throughput: hardware or API region, concurrency, input/output length, and measurement procedure.
- For cost: currency, billing tier, input/output assumptions, caching, and tool charges.
Model pages display source links and collection dates where present. A collection date describes when data was retrieved, not when a benchmark was run. Missing evidence is not filled with invented test results. Until these fields are complete, this catalog cannot support fully reproducible comparisons.
How to use the results
Use the catalog to find candidates. Follow their source links and run a matched evaluation on your own tasks before purchasing or deploying. Different harnesses, budgets, and test populations can produce incomparable scores. An aggregate quality score also does not establish a hallucination rate or safety classification.
Editorial standards · Report a correction · Return to the leaderboard