AIRANKS — The Authoritative Rankings for AI Web Content

AIRANKS measures AI visibility: we ask AI models real product and service questions, capture the complete answers as immutable observations, and publish what they contain — which brands were mentioned, which domains were cited, and which exact pages were linked. Every domain gets an AIR score from 1–10 (a decile of visibility in the active dataset; 0 means insufficient data), with the methodology in the open.

Skip to main content

AIRVER. FIGHTING — 2026

Artificial Intelligence Rankings

how a ranking gets made

A rank without its stabilityis not a finding.

Every number on this site is a measurement, not an opinion — and like any measurement it comes with error bars. This page explains how they're produced, so you can judge them instead of just trusting them.

runs / phrase
100
models
1
resamples
1,000
min occurrences
3
publish floor
25 runs

01 — repetition

Why we ask the same question 100 times.

We send each question to 100 separate model runs, split across 1 models. A single answer is an anecdote. A hundred answers is a distribution.

We deliberately do not run at temperature=0 or pin a seed. That would look more deterministic, but it would be a lie — providers batch requests together on shared hardware, and that batch-invariance means two identical requests can still land in different batches and come back different. Determinism at the API layer is not guaranteed even when you ask for it, so we sample at a realistic temperature and let the variation show up honestly instead of hiding it behind a setting that doesn't actually deliver what it promises.

02 — position

Position is first mention, not sentence order.

position_in_response is the order in which a brand first appears, counting distinct entities — not which sentence it's in. A brand named once in the opening line and then repeated three more times later in the same answer still gets credit for appearing first, once. This survives both flowing prose and numbered lists without needing separate handling for either.

03 — confidence

The bootstrap: 1,000 resamples per rank.

For every phrase, we take the runs we actually collected and resample them — with replacement — 1,000 times, recomputing the ranking fresh each time. That gives us a distribution of ranks instead of a single guess.

100 runs 1,000 resamples percentile intervalstability %

rank_stability_pct is the share of those resamples where a brand held the exact same position it holds in the headline ranking. A brand sitting at 96% stability held its spot in 962 of 1,000 resamples — it's a real result. A brand at 48% stability is telling you the ranking is close to a coin flip once you account for the noise inherent in a hundred samples. Both get shown as a plain rank number by every product we've seen except this one. Published research (arXiv:2606.24381) found the #1 spot in LLM brand rankings frequently flips once you resample at n≈100 — which is exactly the regime most rankings, including ours before resampling, get built in.

That's why every rank on this site carries its confidence band next to it, drawn, not just labelled — and there is one colour, ledger green, not a ramp. Uncertainty lives in two things instead: the band's width is the interval itself, wide when a rank is close to a coin flip and narrow when it isn't, and the marker's waver wobbles with an amplitude sized to how often that rank actually flipped across resamples — a locked #1 sits still, a contested #4 visibly refuses to settle. A rank you can't see the confidence of is a rank someone is asking you to trust for free.

04 — the AIR score

A domain's AIR is a decile, 0 to 10.

AIR — Artificial Intelligence Ranking — is one number for how visible a domain is across the models we sample. It's a true decile: NTILE(10) over how often a domain turns up, so a 7 means the domain sits in the seventh band of everything we've measured, not that it scored 70 out of 100 on a formula we invented.

AIR 8
example Eight of ten bands. Countable at a glance — the tick marks the halfway point.
insufficient data
example Fewer than 3 observations. Deliberately out of band — not a bad score, and never coloured like one.

That zero is the part worth arguing about. Every scoring product we've looked at maps thin evidence onto the bottom of its scale, so "we haven't seen enough to say" and "this is bad" come out looking identical. They aren't the same claim, and a metric that can't tell you which one it's making isn't one you should quote.

05 — when we refuse to answer

Below 25 runs, we show no interval at all.

A bootstrap will happily resample three runs and hand back a confidence interval of 1.0000 to 1.0000 and a stability of 100.00%. Those are the most confident-looking numbers on the page, and they come from a sample that cannot support them. So below 25 runs we publish no interval and no stability figure — not a zero, not a wide band, nothing. The phrase still shows how often each brand was named, because that is a raw count and it stays honest at any sample size.

A blank where a stability percentage should be is therefore deliberate. It means we do not have the data to make that claim yet, and we would rather say so than print a decimal we cannot defend. The threshold is published rather than hidden for the same reason.

what this does not measure

Limitations, stated plainly.

  • 01This is the model's API, not its consumer chat app — and whether they agree depends on where the app got its answer. We query the model directly with the bare question and nothing else. The consumer app can also search the live web, and what it reads changes what it says. Asked "what is the best coffee maker", our sample put Technivorm first in every single run. So did the app — when it cited a review publication. In a different session, where it pulled only retailer listings and prices, it led with a brand that appears in none of our runs. Asked about overbed tables, where it cited no publication at all, only one of its five brands showed up in our data.
  • 02So a mismatch with the public chat app is not us being wrong. It usually means the app was reading a shop and we were reading the model. Both are real answers to different questions, and we only claim to measure one of them. Every ranking here is labelled with the model and interface we actually sampled, and we never claim to speak for the app.
  • 03Training cutoffs are often unverified. A model can only recommend what it has read. Where we have confirmed a cutoff date against the provider's own documentation we show it; where we haven't, the page says cutoff unverified rather than printing a date we can't stand behind. An unverified cutoff means a ranking involving recent products may be explainable by the model simply not knowing they exist yet — not by any real preference.
  • 04Brand detection only knows the brands we've told it about. Brands are matched against a gazetteer of known names and aliases. A real brand that isn't in that list yet — a new entrant, an obscure regional name, a rebrand — is invisible to the count even if a model names it constantly. The gazetteer grows over time, but any given snapshot has a floor it can't see below.
  • 05Some brand names are ordinary words, and we handle that imperfectly. A few brands are named after words models use constantly — a CRM called Capsule, a mattress feature called Pod, a marketing tool called Drip. Left alone, those match any answer mentioning coffee pods or drip coffee, and we measured exactly that: one such name held second place on every coffee ranking before we caught it. Names like these now only count inside their own category. The trade is real in both directions — a genuinely cross-category brand can be undercounted on a category it also belongs to, and that kind of miss is invisible, because a brand we failed to match looks identical to one the model never named. We audit for it; we don't claim to have caught every case.
  • 06Exact-sentence echo attribution does not work, and we measured that directly. We tried to find out which web pages a model's exact phrasing traces back to by searching its own sentences verbatim. Against real stored answers, 16-, 12-, 8-, and 6-word fragments returned zero hits every time; only generic 3–4 word filler ("best overall for most") returned anything, and that's too short to attribute to a source. LLM output is novel composition, not copy-paste, so there is usually nothing to find verbatim anywhere on the web. We still show which domains a model explicitly cites by URL — that signal is real — but "who this sentence echoes" is a claim we no longer make.

If a ranking anywhere on this site doesn't show its stability, its confidence interval, or the dataset it came from — and it isn't marked as being under 25 runs — that's a bug, not a design choice.

Back to the rankings