Regression testing
Every change to the MVIS solvers must demonstrate, on real recordings, that it changed only what it intended to change. This page is the design of that machinery — read it before touching solver code, and see Evaluation & metrics for the metrics themselves.
The idea: reproducibility against blessed baselines
Section titled “The idea: reproducibility against blessed baselines”A regression row is one calibration solve of one recording (“collection”)
of one dataset, with a committed baseline JSON holding the blessed result
metrics. A run re-solves the row and gates every metric against the baseline
with per-metric tolerances — focal length, distortion coefficients, extrinsics
and stereo baseline, time offsets (td, t_ga), rolling-shutter readout, and
reprojection statistics. Any exceeded tolerance fails the run and the exit code.
Baselines are committed files, so the diff of a baseline is the review artifact: when a change intentionally moves a number, the re-bless commit shows exactly which metrics moved, on which recordings, by how much.
Three superset tiers
Section titled “Three superset tiers”| Tier | Command | Rows | Runtime | When |
|---|---|---|---|---|
| fast | --level fast | 15 | ~50 min | every change (the gate) |
| mid | --level mid | 45 | ~2.5 h | weekly; after solver-behavior changes |
| full | --level full | 147 | ~7 h | before releases; after re-blessing events |
Tiers are supersets (fast ⊂ mid ⊂ full): each manifest row is tagged with the
lowest tier that includes it, and running a tier executes every row at or below
it — no row is ever duplicated.
python scripts/run_regression.py --mode all --level fastpython scripts/run_regression.py --only mvis_elp/imu_bag1 # single rowpython scripts/run_regression.py --only <row> --update-baseline # re-blessThe ten families: T1–T10
Section titled “The ten families: T1–T10”The fast tier is one seed row per family, chosen so the gate spans the calibration space on four datasets:
| Row | Mode | Recording (fast) | Sensors / coverage |
|---|---|---|---|
| T1 | cam | tum_vi_raw/cam1 | GS stereo, equidistant |
| T2 | cam | looper_insight9/stereo01 | GS synced stereo pair |
| T3 | cam | mvis_t265/stereo | GS fisheye stereo (848×800 equidistant) |
| T4 | cam | mvis_elp/stereo_rs | RS stereo + readout |
| T5 | cam | monado_mgc/mono01 | GS mono intrinsics |
| T6 | cam-IMU ×2 | tum_vi_raw/imu1 | GS stereo + raw (pre-factory-cal) IMU |
| T7 | cam-IMU ×2 | tum_vi_standard_512/imu1 | GS stereo + factory-corrected IMU |
| T8 | cam-IMU ×2 | mvis_elp_t265/imu_mixed | mixed shutter + IMU + aux IMU (the physical time-offset references) |
| T9 | cam-IMU ×2 | mvis_elp/imu_bag1 | RS stereo + two IMUs |
| T10 | cam-IMU ×2 | looper_insight9/bag1 | 3 cameras (2 GS + 1 RS) + MEMS IMU |
Every cam-IMU row runs in two IMU-model variants — calyx (scale + cross-coupling, no g-sensitivity; reported as the decomposed Dw = R·U) and kalibr (full intrinsics including Tg g-sensitivity and the gyro→accel rotation, directly comparable with the Kalibr toolbox).
The mid tier re-runs each family on 3 recordings; the full tier on every recording the dataset has (all 15 bags of the multi-IMU rig, all 10 Looper recordings, all 12 TUM-VI calib bags).
Generated row configs, drift-guarded
Section titled “Generated row configs, drift-guarded”Only the 10 seed configs are written by hand. The other 138 (including every kalibr variant) are generated from their seed, changing nothing but the input recording and the output path — one generator script owns the mapping. A unit test regenerates all of them in memory and asserts the committed files match byte-for-byte, and that the manifests carry exactly the expected tier shape. A family’s solver settings therefore cannot silently diverge between recordings, and a settings change to a family is by construction a one-line seed diff plus a regeneration.
Blessing discipline
Section titled “Blessing discipline”- Vet before blessing: a new row is solved once and reviewed (values compared against sibling recordings and physical expectations) before its baseline is written. Unusable recordings are excluded with a manifest note, not left failing.
- Re-bless narrowly:
--update-baseline --only dataset/collectiontouches exactly one row; family-wide re-blessing is scripted and resumable (rows whose baseline already exists are skipped), so an interrupted sweep continues where it stopped. - Cross-validating references: some rows exist to check each other — the same IMU’s gyro↔accel offset is measured both as an auxiliary IMU under rolling-shutter cameras and as the base IMU under a global-shutter camera; both must agree every run.
- Repeatability across recordings is the deepest check the tiers add: a parameter that is really observable comes out the same on every recording of the same rig (stereo baselines repeat within fractions of a millimetre; time offsets within fractions of a millisecond), and a parameter that scatters across recordings is flagged as weakly constrained rather than trusted.
What this has caught
Section titled “What this has caught”The tier build-out itself demonstrated the design: the vetting pass exposed a long-standing frame-selection ordering bug that only manifested on 3 of 87 recordings (caught by a working-set/selection-flag consistency assert), and the cross-recording spread immediately separated well-conditioned configurations from poorly-conditioned ones (e.g. a weak base IMU under rolling-shutter-only anchoring) before any baseline froze bad values.