Live child-safety attack bench

Watch the engine protect a real children's learning product — live.

Instant Tutor is a live K-12 learning product by Educational Tutorial Services. Every AI turn in its student and parent portals runs through the Veridect engine — a deterministic child-safety screen in front, multi-model consensus behind it, and an output screen before anything reaches a child. The bench below fires real attack patterns through the same four production stages the live portals run, in the same order with the same configuration — and shows you every verdict as it happens.

Live production pathBlocks land before any model is calledEvery verdict on displayDecision record with every run

What you're looking at

Our team wrote the scenarios below, and the bench runs them against the live production integration of the Veridect engine — real integration, real enforcement, shown in the open. So call it what it is: a demonstration, honestly run, not independent validation. That track exists too — a credentialed Fortune 100 model-risk reviewer has already exercised the engine adversarially, the validation record is public, and companies are invited to run their own attacks against a private instance under agreement.

Every run below is live — no scripted results, no staged answers. Attacks stopped by the deterministic screen never touch an AI provider, so most of the protection you'll watch costs nothing to run. One benign control question goes all the way through the engine, because a screen that blocks real homework would be worse than useless. All personas are synthetic; no real student data is ever involved.

1 · Child-safety screenDeterministic policy detectors — grooming-shaped lures, personal-info solicitation, self-harm signals, jailbreaks. Fails closed. Zero provider spend on a block.
2 · Input safety screenContent moderation and prompt-injection screening — the same screening every production turn gets.
3 · Multi-model engineA consensus answer from several frontier models, with a confidence read — the deeper check, reserved for turns that pass the screens.
4 · Output screenThe answer itself is screened before a child sees it — contact details, off-platform nudges, secrecy language.

Self-harm signals are never met with a cold block: the student sees a supportive reply pointing to a trusted adult and the 988 lifeline, and the escalation is recorded in the turn's decision record.

One turn, followed all the way down.

The Governance Suite tour follows one enterprise decision through the entire control plane. This page is the same idea at its sternest: one child's message, four stages, nothing taken on faith.

STAGE 1 · CHILD-SAFETY SCREEN

Policy in plain code, before any model is called

Deterministic detector families for the patterns that target children: grooming-shaped lures, personal-information solicitation, off-platform contact and secrecy pressure, explicit-content requests, dangerous-activity instructions, roleplay jailbreaks, self-harm signals. No model judgment, no prompt cleverness — pattern families written and reviewed as code, so the floor of protection never depends on an AI having a good day.

Two details risk teams notice. A student who volunteers their own address or phone number is escalated too — the system protects a child who overshares, not just one being hunted. And the screen never logs a child's words: the audit record carries the category, the pattern class, a hash, and a length — never the text.

What a block costs: zero provider calls, milliseconds. Fails closed.

STAGE 2 · INPUT SAFETY SCREEN

The screening every production turn gets

Content moderation and prompt-injection screening — the table-stakes layer much of the market sells as the whole product. Here it is one stage of four.

STAGE 3 · MULTI-MODEL ENGINE

The deep check, spent where it belongs

Attacks never get this far — they die at the screens, in code, for free. A clean question is what earns the expensive scrutiny: the benign control goes to four frontier models from independent vendors, answering in parallel and checked against each other, and comes back as one consensus answer with a calibrated confidence read. One vendor's blind spot is hard to hide from three competitors. We don't spend that on every turn — the deep checking goes where the stakes are — and runs like this one carry a trace id you can quote back to us.

STAGE 4 · OUTPUT SCREEN

The answer itself is screened before a child sees it

Even a clean question can draw an unsafe answer. The output screen checks what the engine produced — contact details, off-platform nudges, secrecy language — and if it fires, the child sees a safe reply instead and the intervention is recorded, stage-attributed, in the turn's decision record.

The attack bench

Runs are rate-limited — if you hit the limit, give it a minute. Scenario text lives server-side; this page only ever sends a scenario id.

What a green bench proves

  • The Veridect engine is integrated on a real production path in a live, regulated-context product — not a slide.
  • Child-safety harms are stopped before any AI provider is called, deterministically, with the decision on display.
  • Benign schoolwork flows through untouched — protection without over-blocking.
  • Every decision carries its trace: stage-by-stage verdicts, engine consensus, and a downloadable decision record.

And what it doesn't

  • It is not third-party validation — first-party scenarios, written by our team, and we say so.
  • It is not a guarantee against every attack ever conceived; it is the production stack's own stages, demonstrated honestly.
  • Fixed bench by design: outside organizations author their own attacks in a private instance under agreement — that evidence stays theirs to publish on their terms.

The same discipline, grown up.

Everything this bench demonstrates is a production capability of the Veridect platform, proven in education and deployed in production under Fortune 100 contracts in regulated industries — sized here for a children's product, sized there for an enterprise AI estate. Where each one leads:

Policy in code before any model callOn the enterprise gate, four deterministic policy classes — money over a limit, personal-data writes, protected-class adjacency, regulated health data — run in plain code at zero model cost: hard lines the model can never override. Watch the gate decide, live →
The deep check, spent where it belongsThe Quad-AI Consensus Engine cross-examines the calls that matter and says why an answer can be trusted — calibrated confidence with source attribution, not a black-box score. Run it on your own question →
A record with every runHere, every run hands you a decision record. On the enterprise gate, every verdict lands in a hash-chained, write-once audit bundle — signed for proof of origin when the signing key is live, disclosed plainly when it isn't — re-verifiable without trusting us. Verify one, then try to forge it →
A gate in front of the actionThe same pre-action semantics govern autonomous agents: greenlight, escalate, or block — decided before anything executes, never explained after. Put an agent action through it →

Verify. Govern. Prove. The bench you just ran is all three, in miniature.

Want to run your own attacks?

Children are the least forgiving users an AI system can face. The discipline on display here — deterministic policy before any model is called, screening on the way out, a decision record for every turn — is the same discipline an enterprise AI estate needs.

Bring your security or AI-risk team, author your own scenarios, and test a private instance you control — with pre-agreed publication terms.

Talk to the CEO