Benchmark

Qwen2.5-7B vs Llama 3.1-8B: Fine-Tuning Compared

46 identical experiments. Different base model, same dataset, same hyperparameters, same hardware. The gap in overfitting behavior wasn’t subtle — and it wasn’t in the direction most people would expect.

In progress
Full write-up coming soon
Head-to-head data table ready. Full analysis dropping soon.
Built on this research

Everything in this article is measured, not promised. See it running:

Talk to NOESIS liveSee the 33x benchmarkBack the research