GR3GARY ROBINSON III

AI & QA ENGINEERING LEADER · VIENNA, VA

Build with ambition.
Release with confidence.

I’m Gary Robinson, III. I help teams turn complex software and emerging AI applications into experiences people can trust.

15+ years of engineering experience Federal + commercial systems Active security clearance

EXPERIENCE ACROSS
TEAMS THAT MATTER

Organizations: Department of Defense, National Guard Bureau, US Army, US Air Force, Department of Homeland Security, TSA, USPS, FEMA, Steampunk, Amazon, UnitedHealthcare, BlueCross BlueShield, Rally Health, Sentral, Clearcover, Flyreel AI, RSA, Symantec, Curb (formerly Taxi Magic), Accenture, GR3 VerifAI.

$ verify --candidate "Gary Robinson III"

✓Leads quality on federal & commercial systems15+ yrs

✓Evaluates LLMs, RAG pipelines & agentsGenAI cert

✓Builds QA practices from scratchproven

✓API coverage on every endpoint200+ tests

✓Section 508 accessibility & DevSecOpshands-on

✓Security clearanceactive

PASSED6 / 6 checks · 0 flaky

ABOUT

Building quality practices from scratch.

I’m a QA engineering leader with 15+ years advancing quality, automation and security across web, mobile and enterprise platforms. I’ve stood up QA practices from zero, led teams on federal and commercial systems, and now focus on AI-native quality: evaluating LLMs, RAG pipelines and agents before they reach users.

I bring AI into everyday test design, data analysis and defect triage — and the discipline to prove that it works.

  • GenAI Certified
  • Active Security Clearance
  • B.S. Information Security, GMU
15+
years leading quality
4
quality gates built & maintained — automation, performance, a11y, continuous monitoring
3
AI agents planning, writing & repairing tests
12
teams audited program-wide for Release-on-Demand readiness under SAFe

01 / WHAT I DO NOW

Quality, rebuilt for the AI era.

Through GR3 VerifAI, I help federal, healthcare and commercial teams find out whether their AI does what they think it does — before their users do.

01

LLM output evaluation

Structured assessment of model responses for accuracy, relevance and consistency — turning “it seems fine” into repeatable results.

02

Hallucination detection

Catching confident-sounding claims that aren’t grounded in source material, so wrong answers are flagged in testing, not production.

03

Prompt & RAG testing

Prompt-testing frameworks and retrieval validation that check every stage of the pipeline, from what’s fetched to what’s finally said.

04

Agentic test generation

AI agents that plan, write and repair test cases on their own. Continuously improving the skills of the AI agents so they continue to get better at each of their intended functions.

SEE THE IDEA

Spot the hallucination.

This is the core of LLM evaluation: break an AI answer into individual claims and check each one against its source. Press run and watch the tests execute.

Illustrative example with made-up content — not output from a real model or client system.

$ llm-eval run --source policy.txt --answer response.txt

sourcePremiums are due within 30 days of the invoice date. Late payments incur a $25 fee.

promptWhen is my premium due?

claims3 extracted from the answer

  1. Your premium is due within 30 days of the invoice date. queued
  2. A $25 fee applies to late payments. queued
  3. You can defer payment up to 90 days on request. queued

RESULT awaiting run

02 / SELECTED IMPACT

Quality you can measure.

From establishing QA practices to testing AI systems, I connect engineering rigor with outcomes that move teams forward.

FEDERAL DELIVERYSTEAMPUNK
7months from setup to launch

Built the team. Established the practice. Delivered.

Led four QA engineers and established the processes and tooling for a flood insurance quote application. Built 200+ API tests and leveraged data-driven testing to improve test coverage by 30%.

QA leadershipPostmanSection 508JMeterKatalonDataDogGoogle AnalyticsGoogle ColabAxe DevTools
MOBILE AUTOMATIONFLYREEL AI
90%of previously manual Android tests automated

Less regression time. More room to build.

Created mobile automation from the ground up, cutting regression testing by two days. Connected test management and automation through a GitHub Actions CI/CD pipeline.

AndroidCI/CDTest automation
UI AUTOMATIONGR3 VERIFAI
50%of UI workflows automated in one month

Personally automated half of the UI workflows in a single month, with little to no business requirements or documentation to work from.

DATA-DRIVEN TESTINGSTEAMPUNK
+30%test coverage

A Google Colab notebook processed thousands of address records to power data-driven tests.

RELEASE READINESSSTEAMPUNK
12teams audited

Identified the gaps and improvements needed to move an organization toward Release-on-Demand.

SECURITYACROSS ROLES
WeeklyOWASP ZAP scans

At Rally Health, maintained the threat model and ran OWASP ZAP scans ahead of every release. Also caught a critical voice-and-location data bug before a major government and military release.

03 / THE EXPERIENCE

Deep roots.
Forward focus.

My career spans healthcare, security, consumer technology and federal systems. Bottom line: I catch what breaks trust before your customers ever see it.

Read the full resume
APR 2026 — PRESENTCURRENT

GR3 VerifAI

Founder, AI Consultant & Principal Quality Architect

AI-focused quality consulting for federal, healthcare and commercial clients: LLM evaluation, hallucination detection, prompt testing and RAG pipeline validation.

Highlights
  • Established an AI-focused QA consultancy delivering LLM evaluation, AI data quality assurance and GenAI testing strategy.
  • Engineered an API test suite with Playwright, reaching 100% coverage of 50+ endpoints in one week.
  • Deployed three AI agents that autonomously plan, generate and repair test cases, improving suite reliability and throughput.
  • Deliver LLM output assessment, hallucination detection, prompt testing frameworks and RAG pipeline validation.

JUL 2023 — MAR 2026

Steampunk

QA Engineering Team Lead · USDA / FEMA National Flood Insurance Program

Led quality engineering for a flood insurance quote application, including automation, accessibility and performance testing.

Highlights
  • Led four engineers and stood up the QA practice, processes and tooling that launched the application in seven months.
  • Built and maintained 200+ data-driven Postman API tests covering every endpoint.
  • Adopted Katalon and Axe DevTools, then trained the team on functional, regression, automation and Section 508 accessibility testing.
  • Owned performance testing: defined coverage, automated nightly and weekly runs, and reported each release.
  • Wrote a Google Colab Jupyter notebook to process thousands of address records, lifting coverage 30%.
  • Audited QA practices across 12 teams to prepare for a Release-on-Demand cadence.

APR 2022 — MAR 2023

Sentral

QA Lead · Sentral.com and Reservation Lookup

Automated web and backend testing with Playwright and Postman; introduced continuous monitoring and real-time alerts with ChecklyHQ.

Highlights
  • Designed and automated test plans for the website and backend.
  • Implemented CI and production monitoring with ChecklyHQ for real-time alerting on test results and stability.
  • Administered Jira (workflows, automation, release management, project migration) and facilitated scrum ceremonies.

AUG 2021 — APR 2022

Clearcover

QA Lead · Car Insurance Platform

Built data-driven acceptance tests and improved requirements traceability, coverage visibility and automation integration.

Highlights
  • Wrote automated, data-driven acceptance tests with Playwright, Jest, Postman and CodeceptJS.
  • Drove root-cause analysis and process improvements; led migration to a new test management tool with requirements traceability and automation integration.

JAN 2021 — AUG 2021

Flyreel AI

QA Director / Scrum Master

Established automation frameworks, mobile coverage, test management and CI/CD to support an AI-powered insurance product.

Highlights
  • Implemented test management, automation frameworks and a GitHub Actions CI/CD pipeline with Jira and Slack integrations.
  • Built Android automation from scratch with TestProject, automating 90% of previously manual tests and cutting regression time by two days.
  • Configured ChecklyHQ API performance tests to surface regressions immediately.

FEB 2015 — JAN 2021

Rally Health

QA Lead / Software Engineer / Security Advocate · UHC Find & Price Care

Supported UHC Find & Price Care with UI and API automation, full-stack verification, threat modeling and security scans ahead of weekly releases.

Highlights
  • Developed UI and API tests in Robot Framework (Python/Selenium) and WebdriverIO (JavaScript); built and maintained Jenkins CI.
  • Maintained the application threat model and ran OWASP ZAP scans ahead of weekly releases.
  • Paired with developers on full-stack verification, validating changes before merge and occasionally fixing bugs.

MAR 2014 — FEB 2015

Amazon

QA Engineer II / QA Lead · Amazon Appstore

Drove test planning, coverage and automation for large-scale Amazon Appstore and Amazon Mobile Android releases.

Highlights
  • Scoped features, planned tests and drove automation across multiple large-scale releases.
  • Delegated within the QA team and reported results to leadership as an advocate for quality at every phase.
Earlier experience

Taxi Magic / Curb
Senior QA Engineer / QA Lead · 2013–2014

RSA, The Security Division of EMC
Principal SW Quality Engineer · 2013

Reality Mobile
SQA Engineer / QA Lead · 2011–2013

Symantec
SQA Engineer / QA Lead, Managed Security Services · 2010–2011

Accenture
Software Engineer — Test Analyst / PL/SQL Developer · 2007–2010

04 / THE TOOLKIT

Hands-on depth.
Leadership perspective.

01

AI & LLM quality

  • LLM evaluation
  • Hallucination detection
  • Prompt engineering & testing
  • RAG validation
  • Generative AI
  • AI agents
02

Automation & engineering

  • Playwright
  • Postman
  • Katalon
  • Robot Framework
  • Selenium
  • Appium
  • WebdriverIO
  • CodeceptJS
  • JMeter
03

Quality & security

  • UI & API automation
  • Performance testing
  • Section 508 accessibility
  • Penetration testing
  • DevSecOps
  • Agile leadership
04

Platforms & languages

  • Python
  • JavaScript
  • SQL / PL-SQL
  • Shell
  • HTML
  • JSON
  • GitHub Actions
  • Jenkins
  • Docker
  • AWS
  • Google Cloud
  • Jira
  • TestRail
  • Splunk
  • Wireshark

FOUNDATION

B.S. Information Technology

Information Security
George Mason University · 2012

CONTINUOUS LEARNING

  • Quality and Safety for LLM ApplicationsDeepLearning.AI · 2026
  • GenAI Essentials CertifiedExamPro · 2025
  • The Data Scientist’s ToolboxJohns Hopkins University · 2016
  • Certified Application TesterMIT / Accenture Solutions Delivery Academy · 2009

05 / WHAT’S NEXT

Your next ambitious project.
Let’s make it reliable.

Exploring opportunities in AI quality, engineering leadership and test automation. Let’s talk about what your team is building.

Get in touch