Define “signal” before comparing models
A news classification, a volatility explanation, and a directional forecast have different evaluation targets. State the source inputs, timestamp, expected output, and review rule. A fluent explanation is not an accuracy measurement.
What to evaluate
- Source fidelity: does the response correctly represent the supplied material, with traceable citations?
- Missing information: does it identify an absent or stale input instead of filling the gap with an invented fact?
- Output validity: does the response match your allowed fields and values, and are those values consistent with the source?
- Operational fit: do observed latency, failure behavior, and cost suit the schedule?
- Change sensitivity: does a model or prompt update alter the behavior on the same evaluation set?
A practical shortlist procedure
Start with models available in your workflow and permitted for the data you use. Verify provider settings and costs at the time of the test. Use the structured-output reference where that API feature is relevant; it addresses formatting, not factual correctness.
Run a small development set, refine the prompt and validation rules, and then evaluate a held-out set once the configuration is fixed. Keep all raw responses and the scoring decisions. This is more informative than ranking provider families by an unnamed internal benchmark.
Try the comparison in a workflow
The Mag 7 Cross-Provider Jury shows how separate model reads can be assembled. Inspect the source inputs and each model prompt before adapting it. Use execution logs to keep the original responses available for review.
For a single-model starting point, use the watchlist briefing walkthrough with no order node. Add a second model only after identifying a specific failure that the second review might help detect.
What this page can establish
We have removed earlier numerical accuracy, cost-superiority, and trading-outcome rankings because they were not supported by a published evaluation. There is no universal winner established here. The comparison procedure and consensus guide provide a way to make a narrower decision from your own evidence.