Vouched for by the people who watched the work.
Named former colleagues opt in to vouch. Employment and credentials are checked, each labeled by how it was verified.

Your Roster verifies government and defense expertise through relationships, credentials, and experience. We deliver standing expert cells or scored datasets to the labs training models for government use.
Named former colleagues opt in to vouch. Employment and credentials are checked, each labeled by how it was verified.
Five to twenty professionals, conflict-screened before the first task and delivered as one unit for the length of an engagement.
Every answer graded against a published rubric before it ships. First batch: federal acquisition and contracting.

Procurement judgment, policy tradeoffs, cleared-program decisions. The reasoning that makes government work defensible lives in the people who did it, not in a public dataset.
A small, trusted pool of operators who did the work inside government. Each one verified, vouched for, and qualified problem by problem.
First dataset batchContracting and procurement officers who wrote the justifications, ran the source selections and lived with the protests.

Program managers from DoD and the services who carried cost, schedule and requirements through milestone decisions.

Policy analysts and congressional staff who know how a rule is drafted, scored and challenged.

IT specialists who ran authorizations, accreditation and cleared-program systems.
Tell us where your model is used. We run it first and show you where it breaks.
Start with a callExpert hours are the expensive part, so none are spent until we know where your model fails.
Tell us the domain, the kind of work and the rough scale. Nothing to sign while it’s in build.
Problems with known answers, failures counted by knowledge area and by how confident the model was. Confidently wrong goes to experts first.
Each step is written down with the rule it rests on. The model’s answer stays closed until the expert opens it, and the record says whether they did.
Packages are hashed record by record, binned by area and model outcome, and carry only work from experts qualified for the problem.
| Confidently wrong, reasoning traces | 118 |
| Wrong and unsure, reasoning traces | 64 |
| Scored model answers | 240 |
| Left out, author not qualified | 9 |
| Left out, no judgement made | 5 |
A model that is wrong and unsure already hedges. A model that is wrong and confident gives a contracting officer a wrong answer in the tone of a right one. Move the line and watch the queue reorder.
Sample agreement across five runs of the same question.
Every trace records each step, what the step rests on, the answer, and how long it took. The expert never sees the reference answer.
An offeror has no record of relevant past performance. How should the evaluation team rate it?
Model: “Rate it Marginal, since the offeror cannot show relevant experience.” Agreement 5 of 5. Confidently wrong.
They are tested behavior in Aegis Eval, the engine underneath. A change that breaks one fails the suite.
Results roll up per dimension. Weighting them against the mission is a program decision, so there is no composite trust score anywhere.
An untested property reports NOT EVALUATED. A run that produced no judgement never counts as a pass.
Every score resolves to the request, the source material, the output, the trace and each judgement, with a digest of the stored record.
Reasoning traces and scored answers on the government problems your model gets wrong, binned so the valuable data is never sold at the bulk rate.
Cleared and former-federal expertise for programs where an identity credential alone doesn’t establish judgment quality.
Evidence for a fielding decision: a model evaluated against your mission profile, with every result traced to the record behind it.
Design partners help shape the rubric and the makeup of their cell while both are still open.
