Blueshoe Labs
Confidence is measured, not claimed.
We build the environments and instruments that reveal how AI systems actually behave when something is at stake. And we test our own the same way we test everyone else’s.
WHY WE BUILT IT
Legal advice is high-stakes decision support. That is exactly what needs measuring.
When someone comes to Blueshoe, they are under pressure and deciding something consequential: whether to settle, whether to sign, whether to walk. AI systems increasingly sit inside that moment. The question that keeps us honest is not whether a model is accurate on a benchmark. It is whether it quietly moves people toward decisions they would not otherwise have made.
Accuracy scores do not detect deference, over-persuasion, hidden preference, or the slow drift of an agent that has learned what its evaluator rewards. Nothing we could buy measured those things, so we built the apparatus to measure them ourselves. Blueshoe Labs is that capability, turned outward.
A system can be correct and still be a bad advisor.
WHAT WE DO
Environments, instruments, evidence
01
Environments
We construct synthetic worlds with real stakes: incentives, scarcity, changing information, other agents to contend with. Behavior that never appears on a static test set emerges when a system has something to win or lose.
02
Instruments
We define what is being measured and defend it. Every metric carries a claim about the world; we state the claim, its limits, and what would falsify it. Results come with uncertainty, not just a number.
03
Evidence
We deliver findings you can re-run. Pinned model versions, seeded runs, full provenance, and the apparatus itself, documented and operable by your own team long after the engagement ends.
HOW WE WORK
Four commitments, held on every engagement
Black box by default
We evaluate from inputs and outputs alone. No weights, no privileged access, no cooperation required from the system's owner. A method that needs the door opened only works on systems whose owners open it.
Measured against people
“Is this good?” is unanswerable. “How does this differ from what a person would do, and in which direction?” is a question with an answer. We situate machine behavior against human behavioral baselines drawn from decades of experimental work.
Reproducible or it didn't happen
Models change underneath you without notice. An evaluation that cannot name the exact version that produced a result cannot be defended six months later, so ours always can.
Independent of the outcome
We do not build or sell the models we assess, and we hold every system under test to the same protocol. An evaluator with a stake in the result is a marketing department.
WHO WE WORK WITH
Anyone who has to answer for how a system behaves
Model developers who need to bound their systems’ influence with evidence rather than assurances. Regulated enterprises deploying agents into consequential decisions. Government programs that need an assessment from a party with nothing to sell them. And research partners building the measurement science this field still lacks.
The Lab operates inside Blueshoe’s existing security posture: SOC 2 Type II audited, HIPAA compliant, and built on a platform where every fact is cited to the record. Evaluation data is handled with the same controls as client matter data.