Audit your AI forecast dataset before calling it a benchmark
An AI consensus dataset is easy to mistake for a benchmark. It has scores, multiple advisor perspectives, forecast horizons, and enough rows to make a chart look convincing. Before asking which forecast performed best,
An AI consensus dataset is easy to mistake for a benchmark. It has scores, multiple advisor perspectives, forecast horizons, and enough rows to make a chart look convincing.
Before asking which forecast performed best, I ask a more basic engineering question: what does one row represent, and what evidence does it actually contain?
Here is a small audit of a public release from iPulse AI, the Open Agentic Investment Research Platform we build at Future Edge Group. You can reproduce it without a production account or private market data.
Start with the unit of observation
On October 4, 2026, I audited all 746 records in our historical consensus snapshot dataset. Each row describes an asset's stored consensus output for a particular snapshot and forecast horizon. It is not an individual advisor forecast, a realized return, or an independent trading experiment.
The audit found:
| Check | Result |
|---|---|
| Records / distinct record IDs | 746 / 746 |
| One-year / five-year records | 373 / 373 |
| Missing values in the six required metadata fields checked below | 0 |
Records with model_count = 1
|
746 |
Records marked not_applicable_not_evaluated
|
746 |
The last two rows matter more than a polished leaderboard. The release records forecasts, but it explicitly does not supply a completed performance evaluation. And its model count tells us to be careful about treating different advisor perspectives as independent models.
This is a bounded metadata audit of one public release. It does not validate its source data, forecast quality, or every field in its schema.
Run the audit before plotting the scores
This Python example uses only the standard library and the public Hugging Face viewer API. It reads the snapshots configuration's train split in pages. Here, train is the dataset split name; it does not establish that a model was trained on these records.
import json
from collections import Counter
from urllib.parse import urlencode
from urllib.request import urlopen
query = {
"dataset": "future-edge-group/ipulse-ai-historical-consensus-snapshots",
"config": "snapshots",
"split": "train",
}
rows, offset = [], 0
while True:
params = {**query, "offset": offset, "length": 100}
url = "https://datasets-server.huggingface.co/rows?" + urlencode(params)
with urlopen(url, timeout=30) as response:
page = json.load(response)
rows.extend(item["row"] for item in page["rows"])
offset += len(page["rows"])
if offset >= page["num_rows_total"]:
break
if not page["rows"]:
raise RuntimeError("Incomplete pagination")
required = [
"record_id", "snapshot_id", "forecast_horizon",
"scoring_completed_at_utc", "source_algorithm_version",
"record_checksum",
]
print("rows:", len(rows))
print("distinct IDs:", len({r["record_id"] for r in rows}))
print("missing metadata:", {
field: sum(r.get(field) in (None, "") for r in rows)
for field in required
})
print("horizons:", Counter(r["forecast_horizon"] for r in rows))
print("model counts:", Counter(r["model_count"] for r in rows))
print("evaluation status:", Counter(
r["evaluation_methodology_version"] for r in rows
))
The API serves the current public release. Record your execution time, repository revision, and downloaded file digest when doing a formal study; today's row count is not a promise about the next release. This example checks that checksum fields exist. It does not recompute or verify them.
Distinct voices do not establish independent evidence
An advisor persona can ask a useful different question: one perspective might emphasize business quality while another emphasizes downside risks. That diversity can improve the review process.
It does not establish statistical independence. Shared models, evidence, prompts, and aggregation rules can create correlated errors. Counting perspectives is therefore a different operation from measuring error correlation against realized outcomes.
For a future comparison, I would declare the experimental unit first: model configuration, advisor configuration, asset-horizon pair, or publication snapshot. Then I would preserve those identities through the evaluation rather than treating every displayed opinion as another independent observation.
A consensus score also needs its own interpretation. Agreement is useful metadata; it is not automatically a calibrated probability that a forecast will be correct.
Keep three layers separate
My preferred design has three explicit layers:
- Recorded output: the forecast, its identity, horizon, versions, and publication provenance.
- Evaluation inputs: observed outcomes, data source, observation window, and adjustment rules.
- Evaluation result: the metric, baseline, eligible population, exclusions, and evaluation version.
For example, a one-year forecast that has not reached its declared endpoint should not silently become a failed forecast or disappear from a table. Record it as unresolved under a stated policy. A shorter interim comparison can be useful, but it answers a different question and needs a different label.
Likewise, a financial-health field that does not apply to an asset category should not be converted to zero merely to simplify a chart. Missing, inapplicable, and measured zero are different states.
Before reporting performance, specify the horizon, available-information cutoff, outcome definition, baseline, duplicate policy, and treatment of unresolved records. Those choices belong in the evaluation contract before seeing the result.
Preserve provenance without overstating it
We also work on Forecast Library and the public Open Forecast Receipt specification and verifier. These address the record-preservation side of the problem: making forecast records portable and inspectable.
A matching digest can help detect a changed artifact. It cannot establish that the forecast was insightful, the inputs were correct, or the record was published before the event. A retrospective receipt needs an explicit retrospective label; stronger timing claims require separate evidence.
That separation is deliberate. A trustworthy research workflow should let a reader verify a recorded claim, inspect its origin, and evaluate its outcome as distinct questions.
For developers working on forecasting or agent evaluation: what do you use as your experimental unit when several advisor personas share one underlying model? I'd be interested in approaches that preserve useful perspectives without inflating the apparent sample size.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes â full credit and traffic to the original publisher.