The same input doesn't always get the same answer
Model outputs vary between runs and between versions. A single spot check tells you how one response looked once — not how the system behaves.
run 2 → “about a week”
run 3 → “5–8 business days”
Evalora runs your AI feature against a repeatable suite of test cases every time you change a prompt or model, then shows you — case by case — what improved, what broke, and what it will cost.
Fictional example — the new prompt reads better and breaks a fact.
Switch suites, turn evaluation criteria on and off, filter to regressions, and open any case to compare outputs side by side. The pass/fail results, deltas and release rule update as you change the criteria.
| Case | Input | Baseline | Time | Est. cost | Candidate | Time | Est. cost | Change |
|---|
AI features don't fail like ordinary code. A prompt tweak or model upgrade that fixes the case you were looking at can quietly change behaviour in cases you weren't. Most teams find out from users.
Model outputs vary between runs and between versions. A single spot check tells you how one response looked once — not how the system behaves.
Teams paste a handful of prompts into a playground, eyeball the results and move on. It's slow, hard to repeat, and nobody can say later exactly what was checked.
Without a fixed test set and a side-by-side comparison, it's easy to release a change that improves tone and breaks facts, or cuts cost and drops key details.
Evalora is being designed around the regression-testing loop engineers already know — applied to AI outputs.
Keep test cases, reference answers and source context in versioned suites. Tag by category, import from files, and grow the set from real failures.
Run a baseline and a candidate over the same suite and see outputs side by side, with per-case quality, response-time and estimated-cost deltas.
Combine deterministic checks (format, length, required phrases) with rubric-based, model-assisted scoring. Choose which criteria gate a release.
Highlight claims that aren't backed by the source context you supplied, so invented details get a second look before they reach users.
Route regressions, flags and low-confidence scores to reviewers. Their verdicts are recorded alongside automated scores and used to calibrate them.
Every run records the prompt, model, dataset version and criteria used. Export a release report showing what was tested, what changed and who signed off.
Define inputs, reference answers and source context. Start with the cases you already check by hand.
suite: support-returns@v3Run the current and candidate versions over the whole suite, recording outputs, timings and token usage.
baseline ↔ candidateSee pass/fail per criterion, regressions, improvements, response time and estimated cost — case by case.
Δ quality · Δ time · Δ costReviewers inspect regressions and flagged answers, confirm or overturn automated scores, and add notes.
queue → verdictShip, hold or iterate — with a report that records the evidence behind the decision.
release-report.pdfOur current design for running evaluations at scale. It is a plan, not a description of a production system, and may change as we build.
Containerised evaluation workers that execute test cases, apply deterministic checks and collect timing and token data. Scales with suite size without managing servers.
Coordinates each run: splitting suites into batches, retrying failed calls, waiting on scoring, and marking the run complete.
Stores versioned test datasets, raw outputs and scored results, so any past run can be inspected and reproduced.
Used for calls to supported foundation models and for model-assisted scoring against rubric criteria.
Region choices, data retention, encryption settings and support for models outside Bedrock are still being decided and will be documented before early access begins.
Automated evaluation makes it practical to check many cases on every change. It does not replace judgment. We treat every automated score as a signal that needs to earn trust against human review.
Length limits, required fields, reply language and exact values can be checked with code. These are cheap, fast and repeatable.
Rubric-based scoring by a model can assess things like faithfulness or tone, but it can be wrong, biased toward certain styles, or inconsistent between runs.
Reviewers label a sample of cases. Comparing their verdicts with automated scores shows where a criterion can be relied on and where it needs a person.
A test suite only covers the cases in it. Release reports list the suite, criteria and reviewer decisions so the evidence has clear limits.
| Criterion | Scored by | Calibration |
|---|---|---|
| Format & language | Code check | Spot-check rules |
| Answers correctly | Reference + model | Human sample each suite version |
| Supported by source | Model-assisted | All flags human-reviewed |
| Follows instructions | Model-assisted | Human sample; re-check on rubric change |
Evalora is being built for product and engineering teams shipping AI assistants, AI features in SaaS products, and internal AI tools — teams who change prompts and models regularly and need a dependable way to know what changed.
We're in development. There are no customers, benchmarks or certifications to show yet, so this site doesn't claim any. We're looking for a small group of early-access teams to shape the product with us.
| Development status | In development — pre-release |
| Product availability | Not yet available; early-access list open |
| Live evaluations on this site | None — the preview uses local demonstration data |
| Infrastructure | AWS architecture planned, not yet in production |
| Company legal name | [Company legal name] |
| Registered address | [Registered address] |
| Company number | [Registration number] |
| Founded | [Year] |
| Founders / team | [Names & roles] |
| Contact | [hello@your-domain] |
Not yet. Evalora is in development. You can join the early-access list below and we'll contact you when design-partner places open.
No. The preview uses fictional, hand-written demonstration data stored in the page. Filtering and toggling criteria recalculate results locally in your browser; nothing is sent anywhere and no model is called.
It will show you how outputs perform against the test cases and criteria you define, and how that changes between versions. Automated scores are indicators to be calibrated against human review — not a guarantee of correctness.
No. Regression testing reduces the risk of releasing changes that break behaviour you've tested for. It can't catch failures your test cases don't cover, and AI outputs can still be wrong in production.
The plan is to support models available through Amazon Bedrock first. Support for other providers and self-hosted models is under consideration; we'll confirm specifics before early access.
From recorded token counts multiplied by per-token rates you configure. Estimates are for comparing versions and will not exactly match your provider's bill.
The planned design stores datasets and results in Amazon S3. Regions, retention, encryption and access controls are still being finalised and will be documented before any customer data is accepted.
Early-access teams get to use pre-release versions, share a representative test suite, and give regular feedback. Terms and any pricing will be agreed individually. [Confirm early-access terms]
We're looking for teams who ship AI features and change prompts or models regularly. Tell us a little about what you're building and how you test it today.