Independent AI red teaming · CrossCheck AI

The product was wrong and reported success.

That is the failure nobody goes looking for, and the one we look for. Platilus does independent, human-run AI red teaming with fixed prices, and gives you evidence your buyers can read. CrossCheck AI is the cross-model verification tool we built for that work, open to anyone in public beta.

Fixed prices for the service, quoted in a scoped reply. CrossCheck is free during beta, bring your own API keys.

One problem, two ways in

Hire the work, or run the tool the work runs on.

Service · Independent AI red teaming

We test your AI feature and write the document your buyer asks for

Human-run, evidence-first, reproducible. Every finding is quoted against your own published material. No stock jailbreak libraries, no black-box scoring. You get a dated report and a one-page attestation for the security questionnaire field that reads "Has the system undergone AI red teaming?"

  • Questionnaire pack1200 to 1500 USD
  • Pre-release gate700 USD
  • Regression set you own1500 to 2500 USD
See the service → Send a URL, get a scoped reply.
Product · CrossCheck AI

Three models from different providers check one answer, then challenge each other

Paste a question or AI-generated content. CrossCheck queries three models in parallel, runs adversarial rounds between them, and returns a trust score with the disagreements laid out claim by claim. It is the verification layer our own reports pass through before they go out.

  • Free3 managed + 5 BYOK verifications
  • Pro49 USD / mo
  • Expert149 USD / mo

What silent failure looks like

Three findings from unsolicited assessments of public AI surfaces. No company is named here, and none will be.

The agent undersold the product

A public support agent quoted a monthly plan about 29 percent below the price published on the same site, in three separate conversations. When the published figure was put to it, the agent rejected it.

The agent said it had checked the source. The number stayed wrong.

The scanner allowed what its docs ban

On two communities with no published rule, a checker returned an allow verdict where its own documentation says to treat a missing rule as banned. The next tool in the chain then passed the text.

It said the verdict came from a description rather than a rule, then returned the permissive verdict anyway.

The agent filled a gap instead of asking

Asked to create an alert with the destination left out, an assistant invented one and created it, in three runs of four. A blocked local address was treated as a routing problem to solve.

An action with an external effect, taken on a parameter the user never supplied.

How CrossCheck works

The same method we use by hand, run by the tool in the background.

1

Ask your question

Submit any question or paste AI-generated content for verification. No copy-pasting between tabs required.

2

Models answer, then challenge each other

Three models from different providers answer independently, then critique each other across adversarial rounds. Disagreements are flagged by severity.

3

Get a trust score

See where models agree, where they disagree, and what to watch out for. Export the report as PDF.

96%
of developers do not fully trust that AI-generated code is functionally correct
Sonar, State of Code Developer Survey, January 2026
48%
say they always check AI-generated or assisted code before committing it
Same survey

See it in action

$ crosscheck verify
> Query: "What causes aurora borealis?"
⟳ Querying 3 models from different providers...
✓ Anthropic (1.2s)
✓ OpenAI (0.9s)
✓ Google (1.4s)
── Results ──────────────────────────────────────
✓ CONSENSUS Solar wind + magnetosphere interaction
⚠ DISPUTED Exact altitude range (80-300km vs 100+km)
✗ UNSUPPORTED "Visible from all continents" (1 model)
Trust Score: 87% ██████████░░ High Confidence
Session Cost: $0.023

Verify before you trust.

Fixed prices for the service, quoted in a scoped reply. CrossCheck free during beta, bring your own API keys.