Define the prediction
Specify the target, horizon and units: position at a future time, event timing or another explicit observable. Avoid treating different prediction tasks as one score.
Research · Evaluation framework
A world model represents an environment and predicts how it may change. Does adding measured position or process history improve that prediction? Our evaluation framework sets out how to test the contribution and report its limits.
The question
Compare the same task with and without the proposed spatial or historical context. Keep model version, scenario and evaluation rules fixed wherever possible.
Specify the target, horizon and units: position at a future time, event timing or another explicit observable. Avoid treating different prediction tasks as one score.
Record model versions, existing inputs, calibration and available compute. Compare against a baseline that receives the same evaluation opportunities.
Add the input under investigation and record timestamps, coordinate frames, uncertainty and missing observations. Examine stale or inaccurate input as well as ideal conditions.
Proposed protocol
Separate development scenarios from held-out evaluation. Record the split and check for information leakage between training, tuning and measurement.
Use repeated runs and documented seeds where applicable. Include sensing loss, delayed data and conditions outside the calibration set; report failures alongside aggregate scores.
Provide sample counts, distributions and uncertainty with the selected error measure. A prediction improvement is not automatically a throughput gain or a safety improvement.
Publication status
This page sets out an evaluation framework. It does not publish a completed benchmark, validated percentage improvement or commissioned field result.
A publishable result should identify the dataset or reproducible scenario, evaluation code, versions, hardware, configurations and material exclusions.
Distinguish measured observations, synthetic data and simulation output. Document collection permission, calibration and limitations on public disclosure.
Make the method and its limits available for review. Where data cannot be shared, state that constraint and describe which parts can still be reproduced.
Start with your operation
Have a dataset, a repeatable scenario or a prediction worth testing? Describe the question and the evidence available, without sharing sensitive data.