Writing Evals for AI Tools
Test automation applied to a new domain. Build your own evaluation framework for AI/LLM systems.
The Inspection Problem
Everyone's adding AI to their products, but how do you test it? Most teams are "vibe-checking" outputs or relying on manual review. That doesn't scale, and it doesn't catch regressions when the underlying model changes.
The Insight
LLM evaluation is essentially test automation applied to a new domain. If you can write test cases, you can write evals. The skills transfer directly. You just need to learn the new patterns for non-deterministic systems.
What You'll Build
In this workshop, participants build their own evaluation framework from scratch:
- Design test cases for non-deterministic LLM-based systems
- Build scoring rubrics (structural assertions, not vibes)
- Handle non-determinism with statistical approaches
- Run real evaluations and interpret results
- Walk away with a framework applicable to any LLM-based feature
Agenda (Half-Day)
- Theory (45 min): What makes AI testing different? Patterns and anti-patterns.
- Demo (30 min): Walk through a real eval framework.
- Task Definitions (45 min): Hands-on design of scoring rubrics.
- Execution (45 min): Generating eval scaffolding, running evals, interpreting results.
- Discussion (25 min): Applying this in your context.
Full-Day Format (Recommended)
The full-day version adds modules on A/B testing protocols, governance gates, and advanced custom scoring design. The extra time allows for deeper exploration and makes the workshop accessible to participants with less testing experience. I recommend the full day whenever scheduling allows.
Who Is This For?
- Testers and QA engineers working with LLM-based tools
- Development teams adding LLM capabilities to products
- Anyone responsible for LLM-based system quality
No AI experience required. Testing experience is the foundation. Participants work in pre-configured cloud environments with AI coding assistants (currently Claude Code), so the focus stays on test design and judgment, not setup. The workshop is tool-agnostic. The patterns apply to any LLM-based system.
Interest Inquiry
This workshop is available for conference tutorials and in-house training.