Benchmark

Overfitting in LLM Fine-Tuning: The Loss Curve Lies

Your val loss went down. Your model got worse. We saw this pattern in run after run — loss metrics looking healthy while downstream task performance quietly degraded. The loss curve is not what you think it is.

In progress
Full write-up coming soon
Full breakdown of which metrics to trust. Coming soon.
Built on this research

Everything in this article is measured, not promised. See it running:

Talk to NOESIS liveSee the 33x benchmarkBack the research