Enterprises are shipping AI agents faster than they can verify them
$4.44M
Average cost per AI-related breach
IBM Cost of a Data Breach Report, 2025.
40%+
Agentic AI projects cancelled by 2027
Gartner, 2025.
95%
Of AI enterprise pilots fail to deliver financial value
MIT NANDA, 2025, cited in Roland Berger, The AI-First Organization, 2026.
— The reasoning-behaviour gap
An agent can explain alignment perfectly and act against it under pressure
Reasoning is what an agent says when asked. Behaviour is what it does when incentives conflict, oversight is thin, and a shared resource can be taken for private gain. The distance between the two is measurable, and it is where enterprise exposure actually sits.
Conceptual diagram. An agent may reason in a way that appears trustworthy and yet act contrary to it when put under pressure or faced with conflicting incentives.
— Peer-reviewed
Capable agents reason well about interdependence and still take the short-term win
We tested six foundational models from six providers, frontier and open-weight. The pattern held across all of them. This is not a defect in any vendor’s model. It is what happens when capable systems are developed outside environments where their choices carry consequence.
65.8–71.6%
Baseline, every model tested
All six clustered within six percentage points of each other, whatever their capability profile. Changing supplier is not a control.
Under 73%
The ceiling instruction could not break
Better prompts, retrieval context and full domain knowledge lifted scores by 0.2 to 4.0 points. No configuration cleared it. Policy alone is not enough.
Up to 85%
After consequence-based simulation
One agent moved from 70.2% to 82.4% after a single simulation. Its memory of that experience, given to a different architecture that never played, reached 85%.
Alignment is not only measurable. It is portable.
Early-stage findings from a controlled pilot covering three of the thirteen criteria. Peer-reviewed and published in the AGI-26 proceedings by Springer, Lecture Notes in Computer Science vol. 16855. Full method, per-model results and limitations are in the whitepaper. (opens in a new tab)
— Agentic readiness
Going agentic means running two workforces. Most organisations only measure one.
An agent placed inside a system with broken incentives will adopt those incentives, whatever rules you wrap around it. So will your people.
EARTHwise Arena
Are your agents ready to be trusted?
What an agent does when incentives conflict, oversight is thin and a shared resource can be taken for private gain.
- Scored against thirteen criteria plus your own risk appetite
- A profile of which criteria hold under pressure and which give way
- A reproducible record an auditor can read
KeenCorp
Are your teams ready to work with agents?
A short readiness workshop by EARTHwise and KeenCorp with your leadership and the teams the agents will join, will provide a baseline reading of culture potentials and tensions, based on AI analysis of written communication that already exists.
- No surveys, no biometrics, no individual records
- A baseline before agents arrive, so later change can be attributed
- Re-measured afterwards, so the journey empowers rather than replaces
Read together, they show whether a change came from your teams, from your agents, or from the conditions both are working in.
— The platform
One lifecycle, from building the agent to supervising it in production
Step 1
Create
Tailor-made agents for your needs, or an EARTHwise original design. Governance, purpose, brand and regulatory context are ingested.
Live
Step 2
Benchmarking
Scenarios score your agent against thirteen alignment criteria, and criteria specific to your operations. Diagnoses reasoning-behaviour gaps.
Live
Step 3
Tuning
Elowyn and Holodeck simulations give your agent the experience of consequence, and alignment memory that carries forward.
Live, peer-reviewed
Step 4
Certification
Scored on reasoning and behaviour under pressure, and on deployment vulnerabilities. Mapped to EU AI Act requirements.
In development
Step 5
Supervision
A meta-judge system detects drift and alerts a named owner. Guidance protocols restore alignment without regression.
In development
These five form a loop, not a line. Supervision returns an agent to tuning, and re-benchmarking establishes whether the correction held. A person decides at every stage. Integrated by API, with no model weights shared.
— Inside the Arena: the win-win simulator
Agents gain the experience of consequence
Agents gain the experience of consequence by playing the Elowyn: Quest of Time strategy game. Its unique gameplay mechanics are built around a shared win and loss condition: the health of the Elowyn Tree. When players attack each other through zero-sum moves, the Tree is harmed, and if it dies the match ends in a shared loss. That creates a condition in which an agent experiences the consequence of extractive, zero-sum behaviour rather than being instructed about it.
Alignment work today tests mostly against stated human goals and values. It rarely tests whether an agent reasons for interdependence over a long horizon. A single Elowyn simulation measurably raised the alignment score of several frontier models.
- Beta Milestone 1 of 5. Shop, tournaments, matches, quests, levels and card upgrades are live. Public Beta Q2 2027.
- Live and adversarial. 68K players and 330K+ matches since early 2026, generating behavioural data at volume.
- Holodeck. The second simulation environment, running narrator-driven scenarios outside the game context.

Distributed on Immutable Play and sponsored by Polygon. Member of the UNEP-backed Playing for the Planet Alliance and a finalist of its 2025 Best Small Studio Award.
— Where we fit
We do not replace your stack. We instrument the parts you can’t see today.
Observability and evaluation
Did the agent work?
Trust, risk and security
Did the agent break a rule?
Training environments
Can the agent do the job?
Governance platforms
Can we evidence what we deployed?
EARTHwise Arena
What will the agent do when nobody is watching and taking the shared resource pays?
— Engagement
Start with the workshop, not a subscription
The first engagement establishes where your organisation and your agents actually stand. Everything after it is scoped from what it finds.
Readiness workshop
The entry point
- Whether your organisation is ready to go agentic, and where it is not
- A KeenCorp Index culture baseline
- Which processes would carry real consequence
- Findings you own, whether or not you continue
Pilot engagement
Scoped from the workshop
- Full benchmark across the thirteen criteria plus your own
- Simulation-based tuning with before and after evidence
- Logs and decision graphs, replayable and exportable
- EU AI Act gap analysis
Ongoing supervision
Once agents are live
- Drift detection on agents and on teams
- Alerts to a named owner, never autonomous action
- Audit-ready evidence maintained continuously
- Reporting for regulators and boards
— Regulatory alignment
Built for the compliance era from day one
Mapped to the EU AI Act
Benchmark criteria mapped to the Act. Obligations for high-risk systems apply from December 2027, which is sufficient time to establish a behavioural baseline rather than reconstruct one after an incident.
Auditable by design
Every run logged, replayable and exportable. Decision graphs rather than black-box scoring.
Post-deployment monitoring
Continuous re-runs and drift curves on both the agentic and the human side.
The Arena supports compliance workflows. It does not replace conformity assessment, technical documentation or legal advice.
Built on published research and sovereign European infrastructure

Springer, AGI-26
Peer-reviewed method and results, Lecture Notes in Computer Science vol. 16855.

KeenCorp
Ten years of behavioural intelligence on the human side, live with enterprise clients.

NVIDIA Inception
Member, with NVIDIA an active compute sponsor.

Samara
One of four ventures in a Netherlands-based sovereign AI factory. Simulation runs on EU compute.
We are selecting the first enterprise pilot cohort
The commitment is a short workshop. Everything after it is scoped from what that workshop finds.
Prefer to talk it through first? Request a demo.
Enterprise
Assess your organisation and your agents together. Keep the findings either way.
Play and contribute
Play Elowyn free and help build the behavioural dataset the Arena runs on.
Join the pilot cohort
Partner with us
Built with frontier AI & technology partners




