Tier one
Test it against something you already know the answer to.
The fastest way to check any of this is to ask the model a question where you know the true answer, then open the trace. We will set that up on the first call.
Fair question. Here is the boundary, drawn plainly, including the parts where you should check our work.
The Determinism Boundary
Every output falls into one of three categories, and we label which is which in the product rather than only on this page.
Sums, ratios, growth rates, scores and rankings are calculated by a deterministic function over data in the model, not produced by a language model. Run it twice on the same data and you get the same answer. You can read the function.
Pulled from a source record and shown with the date it entered the model. No paraphrasing sits between the source and the number.
Summaries, memo narrative, explanations. A language model writes these from computed and retrieved inputs, and every claim carries its source. This is the category where you should read before you sign anything.
In production we label where each figure came from, metric by metric: employee count from one vendor, revenue from another, market growth from a research report in your own document store, financials from your market data subscription. A score is only as checkable as its inputs, so we show the inputs.
When the Model Writes Back
Reading is safe. Writing is where an AI system damages a business, so it escalates rather than acts.
Tier one
Tier two
Tier three
Tier four
The queue learns. Human decisions are cached as embeddings. Next time the evaluating agent meets a similar case it checks how people have decided before, and only escalates something genuinely new. Without that, a system at this scale would hand a person tens of thousands of near-identical decisions and they would stop reading them by the fiftieth.
Write-backs are gated tightly. The graph and your source systems stay in sync in both directions, and the direction that writes into your CRM is the one we are most conservative about.
Tracing an Answer
You see which records it read, when each entered the model, which were valid as of the date you asked about, and what it computed along the way. Each workflow step logs its inputs and outputs, so a screening memo can be reconstructed from its sources without asking us.
In the room
After the fact
In disagreement
Being Wrong on Purpose
A system that never reports a problem is a system hiding one. These are the four we design for.
When two facts disagree
When resolution gets it wrong
When the sources are thin
The part we will not dress up
Where Your Data Sits
Which Model Reads Your Data
The model is configurable. Most clients route to a frontier model under enterprise terms that exclude training on their data, and that is the right default for quality.
The default
When a clause forbids it
Not two systems
Open-weight models trail the frontier on the hardest reasoning. On retrieval, synthesis and drafting against your own material the gap is narrow. We will tell you which of your workflows we think survive the switch and which do not.
Every 60x system is built to operate in regulated environments, with full audit trails, data isolation, and enterprise security standards as defaults, not afterthoughts.
Access and Confidentiality
Every node and every edge in the graph carries its own permission tag, scoped like IAM. Not the folder, not the document: the object. A derived fact inherits a tag the same way a source file does, which matters because derived facts are exactly what a knowledge system produces and exactly what folder-level permissions cannot govern.
Two people asking the same question reach different graphs. Someone without clearance does not get a redacted answer. That part of the graph does not exist for them. Deal teams keep their walls, payroll reaches finance and HR, and every person has a second, private brain that nobody at the firm can read.
Full detail: permissions and the two-brain model. For the temporal side of the audit trail, see what did we know then.
The fastest way to check any of this is to ask the model a question where you know the true answer, then open the trace. We will set that up on the first call.