Back to blog

Token-weighted paired statistics and stricter release gates

Token-weighted paired bootstrap lands across the pipeline, strictness toggles expand, and CI/release pairing expectations become explicit and enforceable.

1 min readInvarLock Team
Matched collections of differently sized weights hang from a common beam, making weighting and pairing part of the comparison.

Release: InvarLock 0.3.3 - Paired bootstrap, strictness toggles, and clearer failures

Highlights

  • Token-weighted paired Δlog-loss bootstrap support (core + primary metric + variance guard).
  • Window pairing enforcement becomes more explicit (overlap/duplicates/mismatch detection).
  • Strictness toggles and report metadata improvements for clearer evaluation outcomes.

0.3.3 tightens the statistical backbone of paired evaluation. The paired Δlog-loss bootstrap work isn’t just a “numbers” change—it’s about making drift conclusions more faithful to what was actually evaluated (token-weighted and paired, not loosely aggregated).

It also makes CI/release expectations blunt and explicit: perfect pairing, non-overlapping windows, and coverage floors aren’t “best effort” anymore—they’re enforced. That’s a theme in this release: fewer fuzzy edges, more things you can confidently point to.

And when things do go wrong, reports carry better context (including evaluation soft-fail metadata), which helps turn failures into something you can diagnose instead of something you just re-run blindly.

Sources

Website documentation explains the maintained workflow. The changes described here belong to InvarLock v0.3.3; the tagged release record preserves that historical scope.

More in Release

Explore nearby related posts.