How small update modules absorb distribution drift without erasing yesterday’s model.
One stream
Many regimes
Bounded updates
The illustrative stream shifts every few hours. Full retraining is too slow; unconstrained online updates are too forgetful.
Solid: incoming target · Dashed: frozen forecast
First term: fit the current window.
Second term: pay for abrupt adapter movement.
lambda: decides how much yesterday constrains today.
Collect the newest labelled slice and drift summary.
Optimize the small module against the current window.
Reject steps that cross the memory budget.
Promote the adapter and keep the base checkpoint intact.
If the per-window gradient and adapter step are bounded, cumulative change grows with the sum of accepted budgets—not the number of observations alone.
This frame communicates proof intuition, not a formal theorem.
Scope: fictional benchmark values for presentation demonstration only.
When should the base itself be allowed to move?
Open the discussion there.