Auctus reads an experiment the way a statistician would if they were not on your team. It answers one question — could this test have detected the effect you are claiming? — and it usually answers no.
Does the assertion “the campaign worked” have proof behind it — or a lift below the smallest difference the experiment could see?
Most campaign wins are smaller than the experiment could see.
A campaign reports +11.9% and gets rolled out. The test ran on 16,000 users against a 2% baseline, which means the smallest lift it could distinguish from noise was +33.4%. The result was checked six times before anyone stopped, which puts the real false-positive rate at 26.5% against the 5% printed on the dashboard. Nothing here is fraud. Every step was reasonable. The number is still not evidence.
Four questions, asked in the same order every time.
The test ran on 16,000. Three critical findings on a result that had already shipped. The mathematics is standard two-proportion power analysis — the contribution is that it is applied to the number before the number is believed, not after it fails to replicate.
The section a competent buyer reads first.
Every report this product emits ends with its own version of this list, generated from the run rather than written by hand. A report cannot be constructed without one — the validator refuses.
Fixed scope, fixed price, and a report you can argue with.
The documents surface, running in your browser — the real engine, installed into the page. Nothing is uploaded, because there is no server to upload it to. Or send one artifact and we will look at it: you get the finding either way, including if the finding is that nothing is wrong.
One product run against your artifacts, with a written report: measures, findings by severity, the resolution floor, and the limits. The price is fixed before the work starts and quoted from the size of your company, not from how the conversation goes.
Where it gets interesting — the findings on one surface routinely explain the numbers on another.
Five surfaces, one decision procedure. Each deploys separately, so one product's failure cannot take another down.
| Product | Surface | In one line |
|---|---|---|
| Nullius | documents | Your retrieval score is measuring your wording. |
| Fiscus | money | Finding the savings is the easy half. Proving one happened is the other. |
| Rima | code | Two checks, chosen because they are high-precision and commonly missed. |
| Arbol | assistant | The layer that answers from your documents, and shows you which ones. |
One test set, one experiment, one statement export, one repository. The first look costs nothing and the finding is yours either way.
rishabh@op2ra.com