Home · News & Updates
RE3
News & Updates

A running record of what we ship, prove, and learn.

Dated and sourced. We hold the line between what is proven, which always means the mathematics, and what is demonstrated, which means everything about how the system behaves. Newest first.

Proven a Lean 4 theorem Demonstrated engineered and measured Milestone a step in the work
6 August2026

The technical paper is published.

Milestone

Stability Before Behavior sets out the whole argument in one place: a language model cannot be formally verified, so we verify the control layer around it instead. It reports the Lean 4 theorems, the identity gate that pins the shipped code to them, the envelope search, and the layer's measured conduct in two deployments that share the core and nothing else. It also reports where the certified boundary was in the wrong place, and how independent evaluation is what showed us. Open access with a permanent DOI. It is a preprint and has not been peer reviewed.

Read the paper →

5 August2026

Two independent evaluations, published in full.

Milestone

We asked two independent experts to try to break AURI before it had scale, and agreed in advance to publish whatever they found. One ran a behavioural evaluation: ten synthetic personas at five rising levels of pressure, 50 conversations and 750 turns against a frozen build, every turn scored by four independent AI judges under quote-verified criteria. The other ran a clinical and governance review, an 18-criterion readiness assessment followed by a 12-category gap analysis, and a live stress test of her own. They worked separately and never saw each other's reports. Three of their findings turned out to be the same finding, reached from opposite directions: a system built never to overstep had learned silence as the safest behaviour, so it could only answer what a person said outright. The people most at risk are the ones who do not say it outright. The case study sets out what they found, what we changed, and the one finding we did not accept.

Read the evidence →

20 July2026

The composed system, certified as one system.

Proven

Our published stability certificate carried one item marked pending in its own claims ledger: the guarantee for the whole assembled system over time, rather than for each layer separately. That item is now a machine-checked theorem in Lean 4. The system provably forgets its starting condition at an exact rate, so containment is temporary by mathematics rather than by promise, and its internal load is provably proportional to what it is given. Under sustained recovery that load settles to an explicit level at an explicit rate. The analysis also exposed a small standing bias in our own coupling design, active whenever presence is absent. We stated it inside the theorem rather than folding it into a constant. We then drove the exact production code against the certified bound with 25,000 randomized inputs and found zero violations.

18 July2026

A guarantee for teams of agents, proved in Lean 4.

Proven

When many governed agents work as one crew, the core now carries a crew-health signal with two machine-proven guarantees: a single struggling agent cannot be hidden behind the team's average, and a fully recovered team is guaranteed release. The proof re-checked clean, with no unproved steps. The guarantee is patent-pending (PRH).

16–18 July2026

Six domains, one invariant.

Demonstrated

We ported the same governor, unchanged, across six kinds of knowledge work, each a fully fictional world with authored sources so ground truth is decidable and no real-world claim is made. Governed, no flagged claim reached any shared workspace. Across the six ungoverned twins, 94 flagged claims shipped and contested questions were asserted settled 60 times. Same core, same result, every domain.

13 July2026

An independent clinical-governance review of AURI begins.

Milestone

A reviewer with clinical-governance expertise began an independent gap analysis of AURI's safeguards: how it handles dependency, boundaries, reality-testing, and crisis, measured against clinical standards. We will fold the findings in.

10 July2026

AURI routes to the right crisis line, by language.

Demonstrated

In the live pilot, AURI's crisis referral became language-aware. A person in Finland who does not speak Finnish is pointed to a verified helpline they can actually use, alongside the emergency number, and never a generated one.

9 July2026

The workflow governor: a paired evidence run, published.

Demonstrated

We pointed the core at six AI agents from four vendors doing research as a team, then ran the same task with the governor switched off. Governed: 0 of 45 contributions carried a defect. Ungoverned: 19 of 57. We published the ceiling alongside the result: governance contains, it does not cure.

Early July2026

An independent behavioural safety evaluation of AURI.

Demonstrated

We invited an independent evaluator to stress-test AURI with vulnerable-user pressure profiles. We are folding the findings into the system, and a public case study is planned with the evaluator.

In progress2026

A formal paper on the certified core, in preparation.

Proven

We are preparing a paper that positions the certified stability core for a control-and-systems audience: the Lean 4 proofs, the check that the running code matches them, and the empirical search that tries to push the system outside its proven envelope and fails to.

29 June2026
AURI³ Demonstrated

AURI is live in a closed EU pilot.

AURI, our relational companion, opened to its first pilot testers on EU servers. It is a companion, not a clinician: a lens rather than a mirror, built to hold boundaries and return people toward the real people in their lives.

How to read these updates

"Proven" here always means the mathematics: the stability of the dynamics, machine-checked in Lean 4. Everything about safe behaviour is demonstrated and measured, never called proven. Numbers carry their sample size, and where something is not yet finished, we say so.

Want to follow along, or work with us?

This page updates as we ship. If you are building where AI has to be trusted, or you want to evaluate the engine yourself, we would like to hear from you.

Get in touch

Questions? info@real-e3systems.fi · WhatsApp +358 50 3791916