# swarmREFLEX-2.1 — model card (EDGE-H1 A24) · SIMULATION

Pre-registration: `docs/edge/EDGE-H1-PROVE.md` §9 A24 (commit f4148e6, before any data generation or training). All of it is
simulation. Calibration numbers come from the **cal-B** seeds and are not evidence. The evidence is the frozen fleet block
`e5r21` (seeds 65100–65115).

## What it is

swarmREFLEX-2.0's recipe, unchanged:
- a per-contract XGBoost chooser plus harm boosters;
- grid max_depth {4, 6, 8} × eta {0.03, 0.08}, up to 4,000 rounds, early stopping after 100;
- Platt scaling on cal-A;
- per-contract thresholds on cal-B, using Clopper–Pearson **and** an episode-cluster bootstrap, upper-95 harm ≤ 0.5 %.

It uses the same 129 observable features and the same declared truth oracle as 2.0. **Only the data changed:**
- more of it;
- from the ±0.5 m UWB world **and** the ±5 m narrowband world (NB2 stand-in, assumed parameters).

In arms E5R21 / E5R21NB it replaces the rule envelope. When it abstains, the rule contract decides, as in E5R.

| | |
|---|---|
| Artifact | `isaac/reflex/edge-r21/out/swarmreflex-2.1.json` (sha256 `885bb590…0376`, model hash `bcebc14a…2792`) |
| Meta | `isaac/reflex/edge-r21/out/swarmreflex-2.1.meta.json` (grid choices, per-world cal-B) |
| Trainer | `train_swarmreflex21.py`, generated by `make_trainer21.py` from the frozen A21 trainer (unchanged); image `swarmcortex-reflex-train:1` (CPU) |
| Trained | 2026-10-04T05:27:32Z, 163 s wall |

## Data

Behaviour policy: the no-LLM twin of E4R, 60 min per run.
- **R1WR**, ±0.5 m: the 500 A21 runs (seeds 66000–66499, raw sha256 `769bc781…7920`, all used as train), plus new even seeds 67000–67798.
- **R1WRNB**, ±5 m NB2: odd seeds 67001–67799.
- Splits on the new seeds: train 67000–67599; cal-A 67600–67699; cal-B 67700–67799.
- 800 new runs, 0 failed, none missing.
- `data/manifest.json`: raw sha256 `6645727…3b77`. The jsonl files are not committed and can be rebuilt by `build_dataset21.py`.

| world · contract | train rows | cal rows | rule action correct (train) |
|---|---:|---:|---:|
| ±0.5 m · blocked-zone-reroute | 160,088 | 19,151 | 10.5 % |
| ±0.5 m · downstream-station-hold | 102,154 | 12,708 | 68.0 % |
| ±0.5 m · charger-slot-power-cap | 44,230 | 5,570 | 60.7 % |
| ±5 m · blocked-zone-reroute | 67,447 | 21,954 | 8.6 % |
| ±5 m · downstream-station-hold | 37,413 | 12,625 | 67.4 % |
| ±5 m · charger-slot-power-cap | 16,177 | 5,428 | 59.7 % |
| **total** | **427,509** (2.0: 153,690) | **77,436** | |

## Calibration and gate (cal-B — calibration, not evidence)

| contract | world | cal-B decisions | admitted | harm events among admitted |
|---|---|---:|---:|---:|
| blocked-zone-reroute | ±0.5 m | 9,518 | 9,462 (99.4 %) | 0 |
| blocked-zone-reroute | ±5 m | 11,092 | 11,046 (99.6 %) | 0 |
| downstream-station-hold | ±0.5 m | 6,397 | 5,345 (83.6 %) | 26 (0.49 %) |
| downstream-station-hold | ±5 m | 6,318 | 5,254 (83.2 %) | 12 (0.23 %) |
| charger-slot-power-cap | both | 5,561 | 0 | — |

Pooled admitted on cal-B (non-fallback) is **80.0 %**. The thresholds are 0.8128 for blocked-zone and 0.8434 for station-hold. Charger has no threshold.

**A24 gate (binding): at least 1 % of envelope decisions admitted on cal-B. Result: PASSED (80.0 %).**

## Known limits, stated in A24 before training

- **Charger:** the oracle's 10-min value never lets a charge beat `defer_charge` above the hard battery floor. This is a horizon-truncation bias of the declared oracle. It is unchanged, so the model never acts on charger decisions and the rule contract decides them.
- **Oracle fidelity (A21 probe):** for blocked-zone and charger, the oracle agrees with real forced reruns at about chance level.
- **Harm bound:** relative to the contract fallback under the oracle, not to realised outcomes.
- **Scope:** simulation research only. Not a safety function.
