Op2ra → the assistant surface
Assistant · árbol — a tree; the structure notes grow into

Arbol

Arbol is the assistant the other four audits make trustworthy. It runs on documents you already own, answers with the notes it used, and — because it is built on Nullius — can tell you whether its own retrieval evaluation means anything.

It makes the assertions, and shows its sources. Then it audits its own evaluation set, and withholds the score when that set cannot support one.

prototype

The situation

The layer that answers from your documents, and shows you which ones.

Every company that buys an internal assistant asks the same question six weeks in: is it actually right? The vendor shows a retrieval score. Nobody audited the test set that produced it. Arbol is the same product with that loop closed — the assistant and the instrument that checks it come from the same codebase, and the instrument was built first.

$ arbol ask "why is the code surface defensive only?" answer ============================================================== It is a decision, not a capability limit. Rima reads a working tree and never touches a running system: no probing, no exploitability testing, and no credential validated against any service. A found key is reported as live and unrotated rather than tested, because testing it is the line. drawn from 3 notes ADR-003 the code surface reads, it does not probe 20260812 Rima — technical deep dive 20260813 product vision — non-goals the same claim appears in the module docstring, the registry, every report's limits section, and this site — checked in CI retrieval over this corpus was audited on 2026-08-13: 59 labelled suites, 38/38 caught, 0/21 false positives

What it checks

Four questions, asked in the same order every time.

  • Grounded answers — every response carries the notes it drew from. An answer with no sources is reported as an answer with no sources, not smoothed over
  • Your corpus, your machine — it indexes a directory of markdown, not a vendor cloud. The corpus never has to leave your network
  • Audited retrieval — the evaluation set is run through Nullius before any score is quoted. That ordering is the whole point
  • Structure that survives — notes carry stable identifiers and explicit links, so the graph is a property of the corpus rather than of the tool reading it

Where it actually is

prototype
not a shipped product
1
corpus it has run on
0
external users
no calibration to quote

This is the one page in the suite with nothing to prove yet. It indexes and answers over a single private vault and has never been run by anyone outside it. Listing it as shipped would be the exact move the other four pages exist to catch, so it is listed as what it is. It becomes a product when a client asks for it — the sequencing rule for this entire suite is that a surface gets built when someone needs it, not before.

What this audit cannot see

The section a competent buyer reads first.

Every report this product emits ends with its own version of this list, generated from the run rather than written by hand. A report cannot be constructed without one — the validator refuses.

  • everything. There is no calibration, no third-party corpus, and no user other than its author
  • the served prototype has no authentication by design and returns full note bodies. It is a local instrument, and the network layer carries the entire security burden
  • whether the assistant helps. The audits measure retrieval; nobody has measured whether answering faster changes a single business outcome

Engagements

Fixed scope, fixed price, and a report you can argue with.

Free

Browser audit

The documents surface, running in your browser — the real engine, installed into the page. Nothing is uploaded, because there is no server to upload it to. Or send one artifact and we will look at it: you get the finding either way, including if the finding is that nothing is wrong.

Fixed

One surface, two to three weeks

One product run against your artifacts, with a written report: measures, findings by severity, the resolution floor, and the limits. The price is fixed before the work starts and quoted from the size of your company, not from how the conversation goes.

Quote

Multiple surfaces

Where it gets interesting — the findings on one surface routinely explain the numbers on another.

The rest of the suite

Five surfaces, one decision procedure. Each deploys separately, so one product's failure cannot take another down.

ProductSurfaceIn one line
NulliusdocumentsYour retrieval score is measuring your wording.
AuctusgrowthMost campaign wins are smaller than the experiment could see.
FiscusmoneyFinding the savings is the easy half. Proving one happened is the other.
RimacodeTwo checks, chosen because they are high-precision and commonly missed.

Send one artifact.

One test set, one experiment, one statement export, one repository. The first look costs nothing and the finding is yours either way.

rishabh@op2ra.com