Methodology before rankings

Satcove Benchmark: Agreement Is Not Accuracy

Satcove compares independent AI answers and reports how they align. That signal is useful, but several models can agree and still be wrong. We will not publish estimated performance bonuses or present model agreement as factual accuracy.

Public benchmark pending

No public Satcove accuracy ranking is claimed on this page today. Results will appear only after the dataset, model versions, execution logs, scoring rules, failures, and comparable denominators can be published together.

What will be measured separately

Panel alignment

How closely the executed model answers align in conclusion and direction. This is what the agreement score measures. It is not accuracy or probability of truth.

Factual accuracy

Whether checkable claims match an answer key or current primary sources, scored independently from model agreement.

Evidence coverage

Whether material claims have relevant sources and whether those sources actually support the conclusion drawn from them.

Freshness

Whether time-sensitive facts were checked against sources current enough for the question and decision date.

Panel integrity

Which vendors and models actually executed, including failures and fallbacks, so nominal panel size is never confused with provider diversity.

Solution quality

Whether the final synthesis addresses the user's constraints, preserves material disagreement, and provides a concrete, proportionate next step.

Planned public studies

The first benchmark releases will focus on decisions where a fluent but unsupported answer can cost time, money or trust.

Purchase decisions

Identical product and price questions sent to six models, with agreement, evidence coverage and recommendation quality scored separately.

Factual verification

Claims with answer keys or primary sources, including freshness and source support — not agreement alone.

Model disagreement

Questions designed to expose differences in assumptions, uncertainty and recommended next steps.

Publication gates

A benchmark becomes decision-useful only when another evaluator can inspect or reproduce it.

  • Versioned dataset and prompts
  • Same questions and denominators for every system
  • Executed model and vendor identities recorded
  • Answer keys or primary-source review for factual tasks
  • Failures, abstentions, and fallbacks included
  • Machine-readable results and methodology published together

What Satcove provides now

A panel of independently generated answers, named convergence and divergence, available sources, and one synthesized recommendation. The agreement score describes the panel; source quality, freshness, and professional review determine how a factual claim should be used.