Research · Evaluation framework

Does more context make a better prediction?

A world model represents an environment and predicts how it may change. Does adding measured position or process history improve that prediction? Our evaluation framework sets out how to test the contribution and report its limits.

Book a demo ↗Explore ↓

The question

Change one input. Compare the prediction.

Compare the same task with and without the proposed spatial or historical context. Keep model version, scenario and evaluation rules fixed wherever possible.

Define the prediction

Specify the target, horizon and units: position at a future time, event timing or another explicit observable. Avoid treating different prediction tasks as one score.

State the baseline

Record model versions, existing inputs, calibration and available compute. Compare against a baseline that receives the same evaluation opportunities.

Isolate the context

Add the input under investigation and record timestamps, coordinate frames, uncertainty and missing observations. Examine stale or inaccurate input as well as ideal conditions.

Next · Method

Proposed protocol

A result should survive repetition.

Keep evaluation separate

Separate development scenarios from held-out evaluation. Record the split and check for information leakage between training, tuning and measurement.

Repeat and challenge

Use repeated runs and documented seeds where applicable. Include sensing loss, delayed data and conditions outside the calibration set; report failures alongside aggregate scores.

Report the right metric

Provide sample counts, distributions and uncertainty with the selected error measure. A prediction improvement is not automatically a throughput gain or a safety improvement.

Next · Publication

Publication status

Method first. Numbers after.

This page sets out an evaluation framework. It does not publish a completed benchmark, validated percentage improvement or commissioned field result.

Reproducibility package

A publishable result should identify the dataset or reproducible scenario, evaluation code, versions, hardware, configurations and material exclusions.

Evidence provenance

Distinguish measured observations, synthetic data and simulation output. Document collection permission, calibration and limitations on public disclosure.

Independent scrutiny

Make the method and its limits available for review. Where data cannot be shared, state that constraint and describe which parts can still be reproduced.

Start with your operation

Bring a question worth testing.

Have a dataset, a repeatable scenario or a prediction worth testing? Describe the question and the evidence available, without sharing sensitive data.

Book a demo ↗ai@systown.ai