<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evals on Fernando Hermida</title><link>https://www.fernandohermida.com/tags/evals/</link><description>Recent content in Evals on Fernando Hermida</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 10 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.fernandohermida.com/tags/evals/index.xml" rel="self" type="application/rss+xml"/><item><title>A Minimal Eval Harness</title><link>https://www.fernandohermida.com/posts/agentic-ai-engineering-journey/08-a-minimal-eval-harness/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://www.fernandohermida.com/posts/agentic-ai-engineering-journey/08-a-minimal-eval-harness/</guid><description>A small, hand-rolled harness for checking whether an agent&amp;rsquo;s output is actually correct, not just well-formed, with a fixed dataset, a scorer per case, and a pass rate.</description></item><item><title>LLM-as-Judge</title><link>https://www.fernandohermida.com/posts/agentic-ai-engineering-journey/09-llm-as-judge/</link><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><guid>https://www.fernandohermida.com/posts/agentic-ai-engineering-journey/09-llm-as-judge/</guid><description>When correctness is subjective, grade agent output with a second, structured LLM call instead of eyeballing every run.</description></item></channel></rss>