Galtea@GalteaAIMay 7We benchmarked 6 Q&A generation frameworks for AI evals using the same model, calibrated judges, and documents. Link to the benchmarks below 👇124
Galtea@GalteaAIApr 30We tested a generic GPT-4.1 LLM judge with a standard faithfulness prompt. 30 examples. 9 hallucinations. Hallucinations caught: 0/9. If your eval pipeline runs this way, it's producing false confidence exactly where confidence is most dangerous. How to optimize your LLM Judge for AI evaluations (And why most teams get it wrong) | Galtea BlogFrom galtea.ai1281