What consensus does
Give several models the same input, collect their structured decisions, and apply a rule to the results. That rule might require unanimity, a majority, or a review on disagreement. The architecture makes differences visible. It does not establish that the agreed answer is correct.
Models can share training material, rely on the same flawed source, or follow the same misleading prompt. Different provider names are not evidence of independent errors. Measure agreement and correctness separately, using recorded inputs and outcomes appropriate to the task.
Inspect a first-party workflow
The Multi-LLM Consensus Trader shows a real template design with model review, a consensus gate, and a paper destination. Open its node graph and use the LLM node documentation to inspect the model settings. This is a configuration example, not a published performance study.
- Freeze the input. Send each model the same source data and timestamps. Record the prompt and model version with the run.
- Define the response. Require a decision from an allowed set, supporting source identifiers, and a missing-data field. Validate both the shape and the content.
- Define disagreement. Specify the minimum valid responses, what counts as agreement, and what to do when a model times out. A missing response should not silently count as approval.
- Keep execution checks separate. Account mode, order constraints, and permitted assets belong in deterministic checks after the model output.
- Inspect the paper run. Read each response, the aggregate decision, and the downstream branch. Use the paper review checklist to cover both action and no-action cases.
Evaluate the task you actually have
For a research summary, measure whether claims match the supplied sources. For schema output, measure valid responses and semantic errors separately. For a market estimate, define the prediction timestamp, horizon, resolution rule, and evaluation dataset before measuring accuracy. Do not compare numbers from different tasks as though they were interchangeable.
Record a single-model baseline and a consensus configuration on the same examples. Track error types, disagreements, latency, and credits. Choose the extra model calls only when the recorded benefit justifies their cost for your task.
Structured output is a format check
OpenAI's Structured Outputs documentation explains schema-constrained responses. A well-formed response can still contain an incorrect price, unsupported explanation, or unsuitable action. Validate values against the input and account rules after validating the format.
Correction to the earlier benchmark claims
An earlier version cited a 78% error reduction and roughly 89% directional accuracy without a linked dataset, methodology, or reproducible evaluation. Those figures have been removed. This article makes no numerical accuracy or investment-performance claim for NickAI consensus.