Skip to content

Regression testing

Every change to the MVIS solvers must demonstrate, on real recordings, that it changed only what it intended to change. This page is the design of that machinery — read it before touching solver code, and see Evaluation & metrics for the metrics themselves.

The idea: reproducibility against blessed baselines

Section titled “The idea: reproducibility against blessed baselines”

A regression row is one calibration solve of one recording (“collection”) of one dataset, with a committed baseline JSON holding the blessed result metrics. A run re-solves the row and gates every metric against the baseline with per-metric tolerances — focal length, distortion coefficients, extrinsics and stereo baseline, time offsets (td, t_ga), rolling-shutter readout, and reprojection statistics. Any exceeded tolerance fails the run and the exit code.

Baselines are committed files, so the diff of a baseline is the review artifact: when a change intentionally moves a number, the re-bless commit shows exactly which metrics moved, on which recordings, by how much.

TierCommandRowsRuntimeWhen
fast--level fast15~50 minevery change (the gate)
mid--level mid45~2.5 hweekly; after solver-behavior changes
full--level full147~7 hbefore releases; after re-blessing events

Tiers are supersets (fast ⊂ mid ⊂ full): each manifest row is tagged with the lowest tier that includes it, and running a tier executes every row at or below it — no row is ever duplicated.

Terminal window
python scripts/run_regression.py --mode all --level fast
python scripts/run_regression.py --only mvis_elp/imu_bag1 # single row
python scripts/run_regression.py --only <row> --update-baseline # re-bless

The fast tier is one seed row per family, chosen so the gate spans the calibration space on four datasets:

RowModeRecording (fast)Sensors / coverage
T1camtum_vi_raw/cam1GS stereo, equidistant
T2camlooper_insight9/stereo01GS synced stereo pair
T3cammvis_t265/stereoGS fisheye stereo (848×800 equidistant)
T4cammvis_elp/stereo_rsRS stereo + readout
T5cammonado_mgc/mono01GS mono intrinsics
T6cam-IMU ×2tum_vi_raw/imu1GS stereo + raw (pre-factory-cal) IMU
T7cam-IMU ×2tum_vi_standard_512/imu1GS stereo + factory-corrected IMU
T8cam-IMU ×2mvis_elp_t265/imu_mixedmixed shutter + IMU + aux IMU (the physical time-offset references)
T9cam-IMU ×2mvis_elp/imu_bag1RS stereo + two IMUs
T10cam-IMU ×2looper_insight9/bag13 cameras (2 GS + 1 RS) + MEMS IMU

Every cam-IMU row runs in two IMU-model variants — calyx (scale + cross-coupling, no g-sensitivity; reported as the decomposed Dw = R·U) and kalibr (full intrinsics including Tg g-sensitivity and the gyro→accel rotation, directly comparable with the Kalibr toolbox).

The mid tier re-runs each family on 3 recordings; the full tier on every recording the dataset has (all 15 bags of the multi-IMU rig, all 10 Looper recordings, all 12 TUM-VI calib bags).

Only the 10 seed configs are written by hand. The other 138 (including every kalibr variant) are generated from their seed, changing nothing but the input recording and the output path — one generator script owns the mapping. A unit test regenerates all of them in memory and asserts the committed files match byte-for-byte, and that the manifests carry exactly the expected tier shape. A family’s solver settings therefore cannot silently diverge between recordings, and a settings change to a family is by construction a one-line seed diff plus a regeneration.

  • Vet before blessing: a new row is solved once and reviewed (values compared against sibling recordings and physical expectations) before its baseline is written. Unusable recordings are excluded with a manifest note, not left failing.
  • Re-bless narrowly: --update-baseline --only dataset/collection touches exactly one row; family-wide re-blessing is scripted and resumable (rows whose baseline already exists are skipped), so an interrupted sweep continues where it stopped.
  • Cross-validating references: some rows exist to check each other — the same IMU’s gyro↔accel offset is measured both as an auxiliary IMU under rolling-shutter cameras and as the base IMU under a global-shutter camera; both must agree every run.
  • Repeatability across recordings is the deepest check the tiers add: a parameter that is really observable comes out the same on every recording of the same rig (stereo baselines repeat within fractions of a millimetre; time offsets within fractions of a millisecond), and a parameter that scatters across recordings is flagged as weakly constrained rather than trusted.

The tier build-out itself demonstrated the design: the vetting pass exposed a long-standing frame-selection ordering bug that only manifested on 3 of 87 recordings (caught by a working-set/selection-flag consistency assert), and the cross-recording spread immediately separated well-conditioned configurations from poorly-conditioned ones (e.g. a weak base IMU under rolling-shutter-only anchoring) before any baseline froze bad values.