FP8 & Precision

Why FP8 Training Silently Destroys Model Quality

We ran the same fine-tuning job twice. Same model, same data, same GPU. One was FP8, one was BF16. The quality difference wasn’t 5%. It wasn’t 10%. We weren’t prepared for what we found.

In progress
Full write-up coming soon
The number is in the paper. We’re still deciding how to frame it.
Built on this research

Everything in this article is measured, not promised. See it running:

Talk to NOESIS liveSee the 33x benchmarkBack the research