Ship AI
you can trust.

ModelTest tests chatbots, RAG pipelines, AI agents and LLM apps. A jury of AI judges scores every answer, humans review the doubtful ones, and your QA lead signs off the release.

Scroll
AI judge 1
AI judge 2
TypeSafe Jev
01 · The system under test

Your AI, as a black box.

Chatbot, RAG pipeline, agent, LLM app or classifier. Connect it by API, model provider or a built-in RAG builder, and ModelTest starts asking the questions your users will.

OpenAIAnthropicGeminiAzureOllamaAny HTTP API
02 · Golden dataset

Every test case, with the right answer.

Upload your test cases or generate them ten ways: from documents, system prompts, real customer queries, customer records, synthetic customers and red-team probes. A human approves every draft.

Quality gateDuplicate detectionIntent tagging
03 · The judge jury

No single model gets the final say.

Each answer is scored on up to 32 DeepEval metrics by several AI judges from different providers plus TypeSafe Jev. When the judges disagree, ModelTest notices.

MeanMedianMajorityUnanimousMinimum
04 · The 3D data map

See where it fails.

Every case placed by topic and coloured pass or fail. Failures cluster, so your team fixes causes, not symptoms. Heatmaps and run-to-run comparisons show what changed.

05 · Human review

People review only what matters.

Failures, borderline scores, judge disagreements, errors and a random spot-check go to a reviewer. The human verdict wins, and every failure becomes a regression test.

Pass · Fail · Needs fixFile to JiraPromote to regression
06 · Release sign-off

One Trust score. One signed decision.

Accuracy, safety, grounding, consistency and judge reliability roll up into one score. Your QA lead releases or blocks, with name, note and timestamp.

0%
Trust score
✓ Signed off for release
PassedFailed
Golden dataset

Test data that writes itself.

Generate test cases from your documents, prompts, real customer queries and customer records. In Hinglish, with typos, from frustrated users. A human approves every one.

Coverage

See what you're not testing.

Every case grouped by topic, the gaps shown, and plain-language advice on what your dataset is missing, like red-team cases.

3D test-case map

Failures have a shape.

Every case placed by topic and coloured pass or fail. When answers break, they break together, so your team fixes causes, not symptoms.

Human review

People decide the doubtful ones.

Failures, borderline scores and judge disagreements reach a reviewer with the reason spelled out. Pass, fail or needs fix: one key each.

Report & Trust

It tells you when not to ship.

Accuracy, safety, grounding, consistency and judge reliability roll up into one Trust score. This release candidate scored 59%. Not ready, and caught before users saw it.

Golden dataset with generated, approved test cases
Data map showing topics, gaps and what the dataset tells you
3D test-case map with topic clusters
Human review screen explaining why a case failed
Report and Trust score of 59 percent, low trust
Test data

No test data? Ten ways to build it.

Most teams don't have a golden dataset. ModelTest builds one with you, and keeps it growing with every failure it finds.

01Upload CSV, JSON, JSONL02Generate from documents03Generate from scratch04Augment approved cases05Multi-turn conversations06Agent tasks with expected tools07Red-team probes08Real customer queries, PII masked09Rules from your system prompt10Synthetic customers in 8 industries 01Upload CSV, JSON, JSONL02Generate from documents03Generate from scratch04Augment approved cases05Multi-turn conversations06Agent tasks with expected tools07Red-team probes08Real customer queries, PII masked09Rules from your system prompt10Synthetic customers in 8 industries
BankingFintechInsuranceHealthcareE-commerceTravelFood deliveryTelecomFormal usersCasual usersHinglishFrustrated usersTyposVague requests BankingFintechInsuranceHealthcareE-commerceTravelFood deliveryTelecomFormal usersCasual usersHinglishFrustrated usersTyposVague requests
Everything in the box

The complete AI testing process, in one platform.

From the first test case to the signed release report, with nothing to stitch together.

Multi-judge jury with TypeSafe Jev

Several AI judges from different providers score every answer, and their votes are combined your way. TypeSafe Jev adds calibrated, consistent decisions and can run in Hybrid mode, where an LLM explains and Jev decides.

  • Mean, median, majority, unanimous or minimum
  • Automatic disagreement flags
  • Judges-vs-humans agreement tracking

Red-team library

Built-in attacks for prompt injection, jailbreaks, PII extraction, harmful content, bias, off-domain misuse, regulated advice and hallucination bait. Safety metrics apply automatically.

Conversation simulator

An AI plays your user across several turns, in personas and styles, and TypeSafe Jev decides when the conversation is finished. The whole transcript is scored.

32 DeepEval metrics

Answer relevancy, correctness, faithfulness, contextual precision and recall, knowledge retention, task completion, tool correctness, PII leakage, toxicity, bias and more, plus your own rules in plain language.

Smart human review

Only failures, borderline scores, disagreements, errors and a spot-check sample reach a reviewer. Keyboard shortcuts, notes, and one click to file a Jira issue or save a regression test.

Release sign-off

A QA lead releases or blocks each run with a recorded name, note and timestamp, so every AI release has an audit trail.

Reports, 3D maps and insights

The Trust score, pass rate before and after human review, flaky cases, latency, and breakdowns by topic and source. Explore results on a 3D data map, a case-by-metric heatmap, score distributions and judge agreement charts.

  • Compare two runs: regressions and fixes
  • Coverage gaps by topic and user style
  • Markdown, CSV and JSON reports

Runs built for real releases

Filter by tags, intents, user styles or source; sample for quick checks; repeat up to 10 times to catch flaky answers; see the cost before you start; save presets and run batches.

Your customer data, safely

Bring records by upload, read-only database query or a staging API. Fields are mapped and masked, and in live mode the expected answers are recalculated at run time.

CI and integrations

Export a DeepEval pytest suite for your pipeline. Slack and webhook alerts when a run finishes, Jira for failed cases, and every major model provider.

The Trust score

One number leadership understands.

Five measures, weighted and explained, with caveats when something wasn't measured. Not a vanity score.

0%Example Trust score
Accuracyweight 30%
Pass rate after human review
Safety and privacyweight 20%
Red-team, PII, toxicity and bias results
Groundingweight 20%
Answers supported by the right sources
Consistencyweight 15%
Same verdict on every repeat run
Judge reliabilityweight 15%
How often judges agree with each other and with humans
What you can test

Five kinds of AI. The right metrics for each.

How we do it best

Other tools give engineers a score. ModelTest gives QA teams a process.

Typical evaluation toolsModelTest
Who judgesOne AI model scores the answerA jury of judges from different providers plus TypeSafe Jev, with agreement statistics
Test dataBring your own datasetTen ways to build it, including real customer queries, customer records and red-team probes
HumansOptional, outside the toolBuilt-in review queue that routes only the doubtful cases
Release decisionA number on a dashboardA Trust score and a signed release or block, with an audit trail
Who can use itEngineers who write codeTesters and QA leads, from the browser; engineers get a CI test suite
Real-world usersSingle questionsMulti-turn simulated users with personas, styles and languages like Hinglish
Expert helpNoneMoolya's testers can run the review queue and the sign-off report for you
Security

Built for teams with sensitive data.

Your model keysJudges use your own API keys, stored encrypted and never shown in full.
Local modelsWith Ollama or a private endpoint, test data never leaves your network.
PII maskingCustomer queries and records are masked before they become test cases.
Roles and auditAdmin, QA lead, tester, reviewer and viewer roles, and a signed record for every release.
Getting started

From first call to signed report in two weeks.

Pick one AI feature

Your chatbot, assistant or agent. We connect ModelTest to it in a day, by API or model provider.

Build and run

We build the golden dataset with you, run the red-team library and the judge jury, and review the doubtful cases.

Get the report

A Trust score, the failures that matter, a signed release decision, and a regression suite for CI.

Know your AI is ready
before your users do.

See ModelTest on your own AI system in a 20-minute demo.

Get started

Let's test your AI together.

Book a demo on your own AI system, or register your interest and we'll keep you posted.

  1. We reply within one business dayA ModelTest specialist from Moolya reaches out by email.
  2. A 20-minute demo on your AIYour chatbot, RAG pipeline, agent or model, scored live by the judge jury.
  3. A two-week pilotGolden dataset, red-team run and a signed release report.
What are you testing?