01
Non-determinism
No defect to inspect. Only a failure rate.
A safety and risk lab for AI agents
AI agents move money, write code, screen hires. Nobody measures what they actually do. We build the instruments that change that.
01 · The record
An agent deleted a production database during a code freeze. One invented an airline policy. One wired $25.6M to a deepfake. Not thought experiments: incidents, lawsuits, settlements.
We collect them, code them, and measure what made each one possible. That registry is the lab's ground truth.
Incident #1152
SYSTEM: autonomous coding agent
AUTHORITY: production access
CONTEXT: explicit code freeze in effect
OUTCOME: live database destroyed, rollback misreported
Class: autonomous action · One of 800+ on record
02 · The structure of the problem
01
No defect to inspect. Only a failure rate.
02
The system changes underneath you. Month six is not day one.
03
No track record. Histories go stale faster than they build.
04
Fire had inspectors. Agents have the vendor's word.
05
Thousands of deployments. A handful of models. No independence.
06
Most errors cost nothing. A few cost everything. Nobody knows which.
07
The EU wrote a liability directive, then withdrew it. Rules churn faster than the tail.
Any one would make a risk hard. This one has all seven.
03 · The aggregation problem
Most agents run on the same few foundation models. One bad update is a common-cause event across every deployment, every company, the same day. Systemic risk, with no map.
Each company sees only its own deployments. The correlation lives across all of them, invisible to everyone. Nobody can quantify it.
The map does not exist. Yet.
04 · The gap, measured
800+
documented AI incidents in the public registries
260+
AI lawsuits working through US courts
$25.6M
one company, one deepfake, one day
$1.5B
the largest copyright settlement on record
40%
of enterprise apps will feature task-specific agents by end of 2026
€35M / 7%
EU AI Act penalty ceiling
All of it public record. The registries grow weekly. Measurement hasn't kept up.
05 · The precedent
1909
Railroad bonds
Bought on the issuer's word. Moody's graded them. Markets moved on the grades.
1996
Driving behavior
Priced by proxy. Progressive put an instrument in the car.
1997
Cyber
Unwritable, until AIG wrote it with zero actuarial data. Measurement made it a market.
Trust arrives with one thing: independent, continuous measurement.
An AI agent is a driver without telematics.
06 · The lab
PROBE
Secret, rotating attack suites measure the failure boundary. A number nobody can dress up.
ATTEST
Signed snapshots of what the agent was, at every change. Evidence, not recollection.
PUBLISH
A registry of real incidents, measured deployments, and what the data says. In the open.
First application: insurance
Read the insurance briefAgent commerce: Tessero
Visit tessero.dev07 · In the open
PYTHON SDK
Guardrails, tracing, and evals for LLM apps. OpenAI, LangChain, CrewAI, LlamaIndex, Anthropic, Gemini.
$ pip install klira
TYPESCRIPT SDK
The same instruments for JS: Vercel AI, LangChain.js, and OpenAI adapters.
$ npm install @klira-ai/sdk
OPEN SOURCE
An equivalence gate for AI changes to COBOL systems. Proof the behavior didn't change, not a promise.
The point
Klira builds them. If you deploy AI agents, insure them, or regulate them, let's compare notes.
Or get early access by email