Flash Cards · Responsible AI

Flash Card: Measuring Bias With Clarify

July 25, 2026 · 1 min read

Generative AI Development · part of The Exam Room

Q. You must measure a GenAI feature for bias and toxicity before launch. Which service?

A. SageMaker Clarify’s foundation-model evaluation (the fmeval library) scores accuracy, toxicity, semantic robustness, and prompt stereotyping (bias). Bedrock model evaluation also offers toxicity and stereotyping metrics.

Why? Bias for a generative model shows up as stereotyping and quality disparity, measured offline, not as label parity.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.