316 models · 51 makers · free to use

Stop guessing which AI model to build on.

AI Windes is a consultant, not a leaderboard. Ask in plain language and get a recommendation, the pricing and benchmark evidence behind it, and the trade-off you’d otherwise discover in production.

Sign in with an email link. No password, no card, nothing to cancel.

316

Models compared

384 catalogue entries

51

Makers covered

frontier labs to open-weights

44%

Open weights

140 self-hostable models

$0.749/M

Median blended price

8 priced at $0

The same catalogue, plotted over time

Frontier capability has risen 9× in three years while the price of a fixed capability level has collapsed. Ten charts, no smoothing.

See AI model trends →

A leaderboard tells you what won. A consultant tells you what to do.

It makes the call

Every answer opens with a recommendation and the score behind it — not a table for you to interpret. Ask for the cheapest model that still handles agentic coding and you get one name, with the competence bar it had to clear.

It shows the evidence

Pricing, context limits, throughput, latency, licence and third-party benchmarks — rendered as charts and tables inline. Every figure is computed from the catalogue; nothing is asserted without a number under it.

It tells you the catch

The runner-up, the capability you give up, the batch endpoint that's half price but asynchronous, and the benchmarks that are simply missing. A recommendation with no downside is a sales pitch.

Three steps, about a minute.

  1. 1

    Sign in with a link

    Enter your email, click the link we send. No password to invent, no card, no trial that quietly ends.

  2. 2

    Describe the job

    “Cheapest model with vision and a 200K context window.” “Compare Claude Opus 5 and GPT-5.2.” “What does 2M requests a month cost?” Plain language is the interface.

  3. 3

    Get a decision

    A named recommendation, the charts that justify it, the trade-off against the runner-up, and an honest note on where the data thins out.

Pricing

It’s free. That’s the whole pricing page.

No plans, no seats, no usage meter, no card on file. Sign in with an email link and ask as many questions as you like.

Try AI Windes

Questions

The ones worth answering before you sign in. Everything else lives in the help docs.

Is it actually free, or free-for-now?

Actually free. No paid tier, no usage cap, no card, no trial that converts. Signing in exists so your chats stay attached to you between visits and so we can tell you when the catalogue changes materially.

If we ever add paid features, nothing you already use starts charging without notice first.

Does it use an LLM to write the answers?

No. Answers are computed directly from the catalogue by a deterministic analyst — no model provider is called and nothing is generated. That buys you three things:

  • The same question always returns the same answer.
  • No figure can be invented; every number traces to a row in the catalogue.
  • It can’t bluff about things it doesn’t know, which is why it declines questions outside its remit instead of improvising.

How does it decide what to recommend?

Your constraints filter the catalogue, then the remaining models are scored on a composite of the benchmarks relevant to your task — coding index for a coding question, Design Arena Elo for a design one, and so on. Weights are renormalised over the metrics each model actually reports.

When you ask for the cheapest model that does something, candidates must clear a competence bar before the price sort runs — otherwise the cheapest answer is a model that technically has the capability flag and can’t do the job. Full method.

What happens when a benchmark is missing?

It’s excluded, never treated as zero. A model with no relevant coverage is left out of the ranking rather than scored badly, and the answer tells you how many models qualified.

This matters more than it sounds: only 213 of 384 entries publish an intelligence index and 117 publish a coding index. Newer and smaller open-weights models are the most likely to be thin. Where the data ends.

Can I trust a recommendation enough to build on it?

Trust it to take you from hundreds of options to two worth evaluating properly — that is the job it does well. Then test those two against your own workload.

Benchmarks measure benchmark performance, which is a proxy for how a model behaves on your task, not a substitute for finding out. Prices are list prices and exclude committed-use discounts. The consultant states these assumptions in every answer rather than burying them.

Why did it refuse to answer my question?

It only answers questions it can ground in the model catalogue. “How do I write Python code to reverse a list” mentions coding but is a request for code, not a model-selection question — so it declines. Rephrased as “which model is best at Python code generation”, it answers.

A confident answer built on no data is worse than a refusal. More on the scope rule.

Where does the catalogue come from?

384 entries covering list pricing, context and output limits, capability flags, measured throughput and latency, serving providers, licensing and knowledge cutoffs — plus third-party benchmark results from Artificial Analysis and Design Arena. We don’t run our own evaluations, take payment for placement, or weight any maker preferentially.

What do you do with my email address?

Send sign-in links and occasional product updates. It’s the only account identifier we hold — there’s no password, no profile, and your chats are never stored on our servers. Ask us and we’ll delete it the same day. Privacy policy.