Benchmark

TruthfulQA: Catches What Other Metrics Miss

MMLU looked fine. Perplexity looked fine. Loss looked fine. TruthfulQA caught what everything else missed — consistently, across different models, different precision levels, different training configs.

In progress
Full write-up coming soon
Benchmark comparison across 88+ runs. Coming soon.
Built on this research

Everything in this article is measured, not promised. See it running:

Talk to NOESIS liveSee the 33x benchmarkBack the research