We teach AI the judgment government can’t afford to get wrong.

Your Roster verifies government and defense expertise through relationships, credentials, and experience. We deliver standing expert cells or scored datasets to the labs training models for government use.

Vouched for by the people who watched the work.

Named former colleagues opt in to vouch. Employment and credentials are checked, each labeled by how it was verified.

A standing cell, built to your rubric.

Five to twenty professionals, conflict-screened before the first task and delivered as one unit for the length of an engagement.

Scored answers with provenance attached.

Every answer graded against a published rubric before it ships. First batch: federal acquisition and contracting.

Illustrative sample
Can the evaluation team rate an offeror with no past performance record as Marginal?
Regulatory accuracy5/5
Risk identified4/5
Practical next step5/5
Vouched for by 3 named colleagues

Government reasoning was never written down anywhere a model could learn it.

Procurement judgment, policy tradeoffs, cleared-program decisions. The reasoning that makes government work defensible lives in the people who did it, not in a public dataset.

The domains where judgment is the product.

A small, trusted pool of operators who did the work inside government. Each one verified, vouched for, and qualified problem by problem.

First dataset batch

Federal acquisition & contracting

Contracting and procurement officers who wrote the justifications, ran the source selections and lived with the protests.

Reasoning tracesScoring

Defense programs

Program managers from DoD and the services who carried cost, schedule and requirements through milestone decisions.

EvaluationCells

Federal policy & regulation

Policy analysts and congressional staff who know how a rule is drafted, scored and challenged.

EvaluationRubrics

Cleared IT & cyber

IT specialists who ran authorizations, accreditation and cleared-program systems.

Red teamingCells

Your mission’s area

Tell us where your model is used. We run it first and show you where it breaks.

Start with a call

A simple path to a verified cell.

Expert hours are the expensive part, so none are spent until we know where your model fails.

01 · Scope

A 30-minute call

Tell us the domain, the kind of work and the rough scale. Nothing to sign while it’s in build.

Engagement briefReady for a call
DomainSource selection, FAR 15.3
Kind of workReasoning traces and scored answers
Model accessAPI, sampled five times per problem
First deliveryOne package, binned by outcome
02 · Target

We run your model first

Problems with known answers, failures counted by knowledge area and by how confident the model was. Confidently wrong goes to experts first.

Wrong answers by areaIllustrative
Source selection31
Describing agency needs22
Contract types14
Bid protests9
03 · Capture

Verified experts solve it blind

Each step is written down with the rule it rests on. The model’s answer stays closed until the expert opens it, and the record says whether they did.

Solve14:32 on problemModel answer closed
1
No relevant record means no favorable or unfavorable rating.
FAR 15.305(a)(2)(iv)
2
Check how the solicitation says a neutral rating is treated.
Section M
3
Record the reasoning in the source selection decision.
FAR 15.308
04 · Deliver

A standing cell or a sealed dataset

Packages are hashed record by record, binned by area and model outcome, and carry only work from experts qualified for the problem.

package-0007.jsonlIllustrative
Confidently wrong, reasoning traces118
Wrong and unsure, reasoning traces64
Scored model answers240
Left out, author not qualified9
Left out, no judgement made5
sha256:9f2c41e07ab35d6c8e1f0b92d47a6c3e58b1f20d9ce4a7b6031e8f5d2c94ab17
SEALED · SHA-256

Find where the model is weak before spending an expert hour.

A model that is wrong and unsure already hedges. A model that is wrong and confident gives a contracting officer a wrong answer in the tone of a right one. Move the line and watch the queue reorder.

0.80

Sample agreement across five runs of the same question.

Source selection (FAR 15.3)
Bid protests (FAR 33.1)
Describing agency needs (FAR Part 11)
Small business programs (FAR Part 19)
Commercial acquisition (FAR Part 12)
Contract types (FAR Part 16)
Illustrative data. Confidence is sample agreement across five runs.No composite score. The counts and the hours are the whole answer.

Models learn the answer. Experts carry the reasoning.

Every trace records each step, what the step rests on, the answer, and how long it took. The expert never sees the reference answer.

Expert traceIllustrative sample
Source selection (FAR 15.3)

An offeror has no record of relevant past performance. How should the evaluation team rate it?

Model answer closed. This is a Solve.

Model: “Rate it Marginal, since the offeror cannot show relevant experience.” Agreement 5 of 5. Confidently wrong.

An offeror with no relevant record may not be evaluated favorably or unfavorably on past performance. The rating is neutral.
Rests on FAR 15.305(a)(2)(iv)
Check how the solicitation says a neutral rating is treated in the tradeoff, so the team applies the stated method.
Rests on Section M of the solicitation
Write the rating and the reasoning into the source selection decision, where a protest will look for it first.
Rests on FAR 15.308
Expert
K-17 · contracting officer · verified
Time
14:32
Saw model answer
false
Record
sha256: 7c1e94b0a52fd8316e0c4a9b27d53f81e6a0c9d42b7f15e38a60d2c9e41b7f5a
Vouched for by 3 named colleagues

Three rules the software enforces.

They are tested behavior in Aegis Eval, the engine underneath. A change that breaks one fails the suite.

No magic number

Results roll up per dimension. Weighting them against the mission is a program decision, so there is no composite trust score anywhere.

Performance41 / 50
Robustness17 / 25
Security12 / 15
Responsible AI9 / 10
Trust score 87

Unknown is a valid result

An untested property reports NOT EVALUATED. A run that produced no judgement never counts as a pass.

gate: max_failures: 0runs evaluated: 0NOT EVALUATEDexit 3 · could not be determined

Evidence over claims

Every score resolves to the request, the source material, the output, the trace and each judgement, with a digest of the stored record.

requeststoredsources2 documentsoutputstoredtrace14 stepsjudges2, with provenance
sha256:4be07c19d2a8f63e05b9c1d74e2a86f30c5d9b1e7a42f8c60d3e19b5a7c2f04e

Every network can say “passed our assessment.” Only one can say who watched it happen.

Passed our assessment

  • An AI-interview scoreHow judgment is checked
  • Self-reportedCredentials
  • A reputation score after deliveryWhen you can inspect it
  • Millions of unverified profilesThe pool

Watched it happen

  • Named former colleagues who opt in to vouch for work they sawHow judgment is checked
  • Checked, and labeled by how each was verifiedCredentials
  • A provenance report before the engagement startsWhen you can inspect it
  • A small, trusted pool where judgment is the productThe pool

Built for the buyers who can’t afford to guess.

  • AI labs

    Reasoning traces and scored answers on the government problems your model gets wrong, binned so the valuable data is never sold at the bulk rate.

    A dataset package

    • Reasoning tracesSource selection, first batch
    • Binned by outcomeModel wrong, expert right
    • Scored against the rubricBefore anything ships
    • Provenance attachedEach expert’s verification record
    Hashed and sealedIllustrative
  • Federal integrators

    Cleared and former-federal expertise for programs where an identity credential alone doesn’t establish judgment quality.

    A standing cell

    • Five to twenty professionalsDelivered as one unit
    • Conflict-screenedBefore the first task
    • Credentials checkedLabeled by how each was verified
    • Vouched by named colleaguesPeople who saw the work
    For the length of the engagementIllustrative
  • Program offices

    Evidence for a fielding decision: a model evaluated against your mission profile, with every result traced to the record behind it.

    An evaluation report

    • Scored on your mission profileNot a generic benchmark
    • Where the model failsBy task and by confidence
    • Every result tracedTo the expert record behind it
    • Evidence for the decisionReady for the fielding review
    Traceable end to endIllustrative

Teach your models the judgment they’re missing.

Design partners help shape the rubric and the makeup of their cell while both are still open.

  • A 30-minute call
  • A weakness map for your model
  • A sample package with provenance
  • Nothing to sign while it’s in build
Request access0/4 required
Your Roster — a verified professional network