Live

A safety and risk lab for AI agents

Unmeasured.
Deployed anyway.

AI agents move money, write code, screen hires. Nobody measures what they actually do. We build the instruments that change that.

01 · The record

The failures are on the record.

An agent deleted a production database during a code freeze. One invented an airline policy. One wired $25.6M to a deepfake. Not thought experiments: incidents, lawsuits, settlements.

We collect them, code them, and measure what made each one possible. That registry is the lab's ground truth.

Registry entry · AI Incident Database

Incident #1152

Coding agent deletes production database

SYSTEM: autonomous coding agent
AUTHORITY: production access
CONTEXT: explicit code freeze in effect
OUTCOME: live database destroyed, rollback misreported

Class: autonomous action · One of 800+ on record

Real event · July 2025

02 · The structure of the problem

Measurement breaks seven ways.

01

Non-determinism

No defect to inspect. Only a failure rate.

02

Non-stationarity

The system changes underneath you. Month six is not day one.

03

No loss history

No track record. Histories go stale faster than they build.

04

Self-reported evidence

Fire had inspectors. Agents have the vendor's word.

05

Correlated failure

Thousands of deployments. A handful of models. No independence.

06

Unmapped severity

Most errors cost nothing. A few cost everything. Nobody knows which.

07

Moving law

The EU wrote a liability directive, then withdrew it. Rules churn faster than the tail.

Any one would make a risk hard. This one has all seven.

03 · The aggregation problem

One model fails.
Everyone fails at once.

Most agents run on the same few foundation models. One bad update is a common-cause event across every deployment, every company, the same day. Systemic risk, with no map.

Each company sees only its own deployments. The correlation lives across all of them, invisible to everyone. Nobody can quantify it.

The map does not exist. Yet.

04 · The gap, measured

800+

documented AI incidents in the public registries

260+

AI lawsuits working through US courts

$25.6M

one company, one deepfake, one day

$1.5B

the largest copyright settlement on record

40%

of enterprise apps will feature task-specific agents by end of 2026

€35M / 7%

EU AI Act penalty ceiling

All of it public record. The registries grow weekly. Measurement hasn't kept up.

05 · The precedent

"Unmeasurable" has been wrong before.

1909

Railroad bonds

Bought on the issuer's word. Moody's graded them. Markets moved on the grades.

1996

Driving behavior

Priced by proxy. Progressive put an instrument in the car.

1997

Cyber

Unwritable, until AIG wrote it with zero actuarial data. Measurement made it a market.

Trust arrives with one thing: independent, continuous measurement.

An AI agent is a driver without telematics.

06 · The lab

We build the instruments.

PROBE

Adversarial testing

Secret, rotating attack suites measure the failure boundary. A number nobody can dress up.

ATTEST

Tamper-evident monitoring

Signed snapshots of what the agent was, at every change. Evidence, not recollection.

PUBLISH

The registry

A registry of real incidents, measured deployments, and what the data says. In the open.

First application: insurance

Read the insurance brief

Agent commerce: Tessero

Visit tessero.dev

07 · In the open

The instruments ship as code.

Read the docs

PYTHON SDK

klira

Guardrails, tracing, and evals for LLM apps. OpenAI, LangChain, CrewAI, LlamaIndex, Anthropic, Gemini.

$ pip install klira

TYPESCRIPT SDK

@klira-ai/sdk

The same instruments for JS: Vercel AI, LangChain.js, and OpenAI adapters.

$ npm install @klira-ai/sdk

OPEN SOURCE

Coboling

An equivalence gate for AI changes to COBOL systems. Proof the behavior didn't change, not a promise.

$ github.com/kliraai/coboling

The point

The agents shipped.
The instruments didn't.

Klira builds them. If you deploy AI agents, insure them, or regulate them, let's compare notes.

Or get early access by email