Theory

What Is Stochastic Sparse LoRA and Why It Works

What if the gradient mask was the regularizer all along? KSS-LoRA applies Bernoulli sparsity directly to LoRA update matrices. The math is simple. The effect on overfitting is not.

In progress
Full write-up coming soon
Full derivation and ablations coming. This one needs to be written carefully.
Built on this research

Everything in this article is measured, not promised. See it running:

Talk to NOESIS liveSee the 33x benchmarkBack the research