LLM output evaluation
Structured assessment of model responses for accuracy, relevance and consistency — turning “it seems fine” into repeatable results.
AI & QA ENGINEERING LEADER · VIENNA, VA
I’m Gary Robinson, III. I help teams turn complex software and emerging AI applications into experiences people can trust.
EXPERIENCE ACROSS
TEAMS THAT MATTER
Organizations: Department of Defense, National Guard Bureau, US Army, US Air Force, Department of Homeland Security, TSA, USPS, FEMA, Steampunk, Amazon, UnitedHealthcare, BlueCross BlueShield, Rally Health, Sentral, Clearcover, Flyreel AI, RSA, Symantec, Curb (formerly Taxi Magic), Accenture, GR3 VerifAI.
$ verify --candidate "Gary Robinson III"
✓Leads quality on federal & commercial systems15+ yrs
✓Evaluates LLMs, RAG pipelines & agentsGenAI cert
✓Builds QA practices from scratchproven
✓API coverage on every endpoint200+ tests
✓Section 508 accessibility & DevSecOpshands-on
✓Security clearanceactive
PASSED6 / 6 checks · 0 flaky
ABOUT
I’m a QA engineering leader with 15+ years advancing quality, automation and security across web, mobile and enterprise platforms. I’ve stood up QA practices from zero, led teams on federal and commercial systems, and now focus on AI-native quality: evaluating LLMs, RAG pipelines and agents before they reach users.
I bring AI into everyday test design, data analysis and defect triage — and the discipline to prove that it works.
01 / WHAT I DO NOW
Through GR3 VerifAI, I help federal, healthcare and commercial teams find out whether their AI does what they think it does — before their users do.
Structured assessment of model responses for accuracy, relevance and consistency — turning “it seems fine” into repeatable results.
Catching confident-sounding claims that aren’t grounded in source material, so wrong answers are flagged in testing, not production.
Prompt-testing frameworks and retrieval validation that check every stage of the pipeline, from what’s fetched to what’s finally said.
AI agents that plan, write and repair test cases on their own. Continuously improving the skills of the AI agents so they continue to get better at each of their intended functions.
SEE THE IDEA
This is the core of LLM evaluation: break an AI answer into individual claims and check each one against its source. Press run and watch the tests execute.
Illustrative example with made-up content — not output from a real model or client system.
$ llm-eval run --source policy.txt --answer response.txt
RESULT awaiting run
02 / SELECTED IMPACT
From establishing QA practices to testing AI systems, I connect engineering rigor with outcomes that move teams forward.
Engineered a Playwright test suite that reached full coverage of 50+ endpoints. Deployed three AI agents to plan, generate and repair test cases.
Led four QA engineers and established the processes and tooling for a flood insurance quote application. Built 200+ API tests and leveraged data-driven testing to improve test coverage by 30%.
Created mobile automation from the ground up, cutting regression testing by two days. Connected test management and automation through a GitHub Actions CI/CD pipeline.
Personally automated half of the UI workflows in a single month, with little to no business requirements or documentation to work from.
A Google Colab notebook processed thousands of address records to power data-driven tests.
Identified the gaps and improvements needed to move an organization toward Release-on-Demand.
At Rally Health, maintained the threat model and ran OWASP ZAP scans ahead of every release. Also caught a critical voice-and-location data bug before a major government and military release.
03 / THE EXPERIENCE
My career spans healthcare, security, consumer technology and federal systems. Bottom line: I catch what breaks trust before your customers ever see it.
Read the full resumeAI-focused quality consulting for federal, healthcare and commercial clients: LLM evaluation, hallucination detection, prompt testing and RAG pipeline validation.
Led quality engineering for a flood insurance quote application, including automation, accessibility and performance testing.
Automated web and backend testing with Playwright and Postman; introduced continuous monitoring and real-time alerts with ChecklyHQ.
Built data-driven acceptance tests and improved requirements traceability, coverage visibility and automation integration.
Established automation frameworks, mobile coverage, test management and CI/CD to support an AI-powered insurance product.
Supported UHC Find & Price Care with UI and API automation, full-stack verification, threat modeling and security scans ahead of weekly releases.
Drove test planning, coverage and automation for large-scale Amazon Appstore and Amazon Mobile Android releases.
Taxi Magic / Curb
Senior QA Engineer / QA Lead · 2013–2014
RSA, The Security Division of EMC
Principal SW Quality Engineer · 2013
Reality Mobile
SQA Engineer / QA Lead · 2011–2013
Symantec
SQA Engineer / QA Lead, Managed Security Services · 2010–2011
Accenture
Software Engineer — Test Analyst / PL/SQL Developer · 2007–2010
04 / THE TOOLKIT
FOUNDATION
Information Security
George Mason University · 2012
CONTINUOUS LEARNING
05 / WHAT’S NEXT
Exploring opportunities in AI quality, engineering leadership and test automation. Let’s talk about what your team is building.
Get in touch