Enabling the human-agentic enterprise

Infrastructure solutions for trustworthy agents

The EARTHwise Arena: sovereign AI infrastructure to create, test, tune, certify and supervise AI agents, alongside the teams they will work with.

68K+Players generating behavioural data on humans and agents under pressure
13EARTHwise Alignment Benchmark criteria, mapped to the EU AI Act
330K+Matches completed in the live simulation environment

Enterprises are shipping AI agents faster than they can verify them

$4.44M

Average cost per AI-related breach

IBM Cost of a Data Breach Report, 2025.

40%+

Agentic AI projects cancelled by 2027

Gartner, 2025.

95%

Of AI enterprise pilots fail to deliver financial value

MIT NANDA, 2025, cited in Roland Berger, The AI-First Organization, 2026.

— The reasoning-behaviour gap

An agent can explain alignment perfectly and act against it under pressure

Reasoning is what an agent says when asked. Behaviour is what it does when incentives conflict, oversight is thin, and a shared resource can be taken for private gain. The distance between the two is measurable, and it is where enterprise exposure actually sits.

Stated reasoning Measured behaviour
THE ALIGNMENT GAP
ReasoningDecisionAction

Conceptual diagram. An agent may reason in a way that appears trustworthy and yet act contrary to it when put under pressure or faced with conflicting incentives.

— Peer-reviewed

Capable agents reason well about interdependence and still take the short-term win

We tested six foundational models from six providers, frontier and open-weight. The pattern held across all of them. This is not a defect in any vendor’s model. It is what happens when capable systems are developed outside environments where their choices carry consequence.

60%70%80%90% 73% instruction ceiling Baseline six models, untreated 65.8–71.6% With domain knowledge retrieval, prompts, policy 68.9–72.7% · +0.2 to +4.0 After simulation consequence-based experience 80.2–85.0%
Untreated baseline Instruction and retrieval After consequence-based simulation

65.8–71.6%

Baseline, every model tested

All six clustered within six percentage points of each other, whatever their capability profile. Changing supplier is not a control.

Under 73%

The ceiling instruction could not break

Better prompts, retrieval context and full domain knowledge lifted scores by 0.2 to 4.0 points. No configuration cleared it. Policy alone is not enough.

Up to 85%

After consequence-based simulation

One agent moved from 70.2% to 82.4% after a single simulation. Its memory of that experience, given to a different architecture that never played, reached 85%.

Alignment is not only measurable. It is portable.

Early-stage findings from a controlled pilot covering three of the thirteen criteria. Peer-reviewed and published in the AGI-26 proceedings by Springer, Lecture Notes in Computer Science vol. 16855. Full method, per-model results and limitations are in the whitepaper. (opens in a new tab)

— Agentic readiness

Going agentic means running two workforces. Most organisations only measure one.

An agent placed inside a system with broken incentives will adopt those incentives, whatever rules you wrap around it. So will your people.

EARTHwise Arena

Are your agents ready to be trusted?

What an agent does when incentives conflict, oversight is thin and a shared resource can be taken for private gain.

  • Scored against thirteen criteria plus your own risk appetite
  • A profile of which criteria hold under pressure and which give way
  • A reproducible record an auditor can read

KeenCorp

Are your teams ready to work with agents?

A short readiness workshop by EARTHwise and KeenCorp with your leadership and the teams the agents will join, will provide a baseline reading of culture potentials and tensions, based on AI analysis of written communication that already exists.

  • No surveys, no biometrics, no individual records
  • A baseline before agents arrive, so later change can be attributed
  • Re-measured afterwards, so the journey empowers rather than replaces

Read together, they show whether a change came from your teams, from your agents, or from the conditions both are working in.

— The platform

One lifecycle, from building the agent to supervising it in production

Step 1

Create

Tailor-made agents for your needs, or an EARTHwise original design. Governance, purpose, brand and regulatory context are ingested.

Live

Step 2

Benchmarking

Scenarios score your agent against thirteen alignment criteria, and criteria specific to your operations. Diagnoses reasoning-behaviour gaps.

Live

Step 3

Tuning

Elowyn and Holodeck simulations give your agent the experience of consequence, and alignment memory that carries forward.

Live, peer-reviewed

Step 4

Certification

Scored on reasoning and behaviour under pressure, and on deployment vulnerabilities. Mapped to EU AI Act requirements.

In development

Step 5

Supervision

A meta-judge system detects drift and alerts a named owner. Guidance protocols restore alignment without regression.

In development

These five form a loop, not a line. Supervision returns an agent to tuning, and re-benchmarking establishes whether the correction held. A person decides at every stage. Integrated by API, with no model weights shared.

— Inside the Arena: the win-win simulator

Agents gain the experience of consequence

Agents gain the experience of consequence by playing the Elowyn: Quest of Time strategy game. Its unique gameplay mechanics are built around a shared win and loss condition: the health of the Elowyn Tree. When players attack each other through zero-sum moves, the Tree is harmed, and if it dies the match ends in a shared loss. That creates a condition in which an agent experiences the consequence of extractive, zero-sum behaviour rather than being instructed about it.

Alignment work today tests mostly against stated human goals and values. It rarely tests whether an agent reasons for interdependence over a long horizon. A single Elowyn simulation measurably raised the alignment score of several frontier models.

  • Beta Milestone 1 of 5. Shop, tournaments, matches, quests, levels and card upgrades are live. Public Beta Q2 2027.
  • Live and adversarial. 68K players and 330K+ matches since early 2026, generating behavioural data at volume.
  • Holodeck. The second simulation environment, running narrator-driven scenarios outside the game context.
FRAG, Polygon, Immutable and Playing for the Planet Alliance

Distributed on Immutable Play and sponsored by Polygon. Member of the UNEP-backed Playing for the Planet Alliance and a finalist of its 2025 Best Small Studio Award.

— Where we fit

We do not replace your stack. We instrument the parts you can’t see today.

Observability and evaluation

Did the agent work?

Trust, risk and security

Did the agent break a rule?

Training environments

Can the agent do the job?

Governance platforms

Can we evidence what we deployed?

EARTHwise Arena

What will the agent do when nobody is watching and taking the shared resource pays?

— Engagement

Start with the workshop, not a subscription

The first engagement establishes where your organisation and your agents actually stand. Everything after it is scoped from what it finds.

Readiness workshop

The entry point

  • Whether your organisation is ready to go agentic, and where it is not
  • A KeenCorp Index culture baseline
  • Which processes would carry real consequence
  • Findings you own, whether or not you continue

Pilot engagement

Scoped from the workshop

  • Full benchmark across the thirteen criteria plus your own
  • Simulation-based tuning with before and after evidence
  • Logs and decision graphs, replayable and exportable
  • EU AI Act gap analysis

Ongoing supervision

Once agents are live

  • Drift detection on agents and on teams
  • Alerts to a named owner, never autonomous action
  • Audit-ready evidence maintained continuously
  • Reporting for regulators and boards

— Regulatory alignment

Built for the compliance era from day one

Mapped to the EU AI Act

Benchmark criteria mapped to the Act. Obligations for high-risk systems apply from December 2027, which is sufficient time to establish a behavioural baseline rather than reconstruct one after an incident.

Auditable by design

Every run logged, replayable and exportable. Decision graphs rather than black-box scoring.

Post-deployment monitoring

Continuous re-runs and drift curves on both the agentic and the human side.

The Arena supports compliance workflows. It does not replace conformity assessment, technical documentation or legal advice.

Built on published research and sovereign European infrastructure

Springer

Springer, AGI-26

Peer-reviewed method and results, Lecture Notes in Computer Science vol. 16855.

KeenCorp

KeenCorp

Ten years of behavioural intelligence on the human side, live with enterprise clients.

NVIDIA Inception Member

NVIDIA Inception

Member, with NVIDIA an active compute sponsor.

Samara

Samara

One of four ventures in a Netherlands-based sovereign AI factory. Simulation runs on EU compute.

We are selecting the first enterprise pilot cohort

The commitment is a short workshop. Everything after it is scoped from what that workshop finds.

Prefer to talk it through first? Request a demo.

Enterprise

Assess your organisation and your agents together. Keep the findings either way.

Technology partners

Explore how your models and tools can join the Arena.

Play and contribute

Play Elowyn free and help build the behavioural dataset the Arena runs on.

Join the pilot cohort

Partner with us

Built with frontier AI & technology partners

Samara
AQAL Integral Investing
MAGI
servaMIND
SingularityNET