Contact Sales

Customer Story

How Khan Academy Strengthens K–12 AI Safety Evaluation with HumanSignal Data Services

In Conversation With

Alex Berry

Senior Technical Program Manager

Pratap Vardhan

Staff AI Engineer

2 × faster delivery
4 million learners around the globe benefit from AI moderation

Custom-built safety for young learners

Khan Academy is a global educational platform that provides mastery learning for more than 100 million students and independent learners. Its AI-powered moderation layer cuts across the entire portfolio of products, including Khanmigo, an AI tutor that provides on-demand support for 4 million teachers, students and parents.

When you build AI experiences for young people, content safety requires especially careful standards and review.

Khan Academy had been running its message moderation through a commercial service that was performing well at industry standard.  The team wanted to take its moderation system further with a custom in-house classifier calibrated specifically to K-12 contexts, giving them direct implementation of trusted moderation policies and thresholds.  To establish ground truth for their in-house model, they needed annotators who understood what "safe for a fourth grader" means in practice.

Bring in the teachers

Khan Academy brought in HumanSignal Data Services to build that ground truth from scratch, working from de-identified message datasets.  Khan Academy had full visibility into Label Studio’s process as it happened and control over use of the datasets.

Every reviewer in the annotator pool came from a background suited to the material: K–12 educators and childcare professionals with degrees in child psychology or human development. Evaluating whether a message is appropriate for a young learner required real pedagogical and developmental context , and that expertise came from people with direct child development experience.

The project ran in two phases:

Phase 1: Classify messages as appropriate or not appropriate for a K–12 setting

Phase 2: Flag credible real-world safety risks for escalation

Every message went to three annotators. Agreement across all three annotators lands at 90%+ and remaining cases are resolved through majority consensus. Any edge cases with lower agreement were flagged for additional human review.

Two Phase Annotation Process for Khan
The first round classifies and the second handles escalations

Results

  • 2× faster delivery: Khan Academy's AI engineering team estimated the same work would have taken twice as much time internally with a comparable team.
  • Faster annotator ramp-up: Internal projects typically require 2–3 calibration rounds after rubric creation. With HumanSignal, annotators reached stable agreement after approximately one round.
  • Low rework: Of all completed tasks, about 5% required additional review, while only 2% ultimately required relabeling after disagreement resolution.

Now Khan Academy has a scorecard for how their upgraded system performs, and a foundation for future moderation model development. Their team can do prompt engineering at scale, and validate future policy changes against empirical benchmarks.

Now Khan Academy has a scorecard for how their replacement system performs, and it's training material for whatever comes next. Their team can do prompt engineering at scale, and validate future policy changes against real data instead of best guesses.

Some of our most critical external communications globally depend on this work, so HumanSignal's time and attention here is helping keep kids safe.

The partnership with HumanSignal Data Services gives Khan Academy a stronger evidence base for evaluating and improving its moderation systems, so the team can stay focused on helping students learn and supporting its mission to provide a free, world-class education to anyone, anywhere.

Need a golden dataset that doesn’t exist yet? Talk to us about HumanSignal Data Services today.