META OPENENV HACKATHON 2026 · ROUND 2 FINALIST

We taught a 1B model to read minds at the negotiation table.

A multi-agent RL environment where AI lawyers face off, predict each other's thoughts, and get judged by a tribunal.

Episodes played: 2,847 Deals struck: 1,914 Walkaways: 933
THE GAP

Real negotiations have hidden information.
Most AI training environments don't.

A negotiation isn't just a sequence of moves — it's a game of beliefs. What do they want? What's their walkaway? What will they accept if I push here?

We built an environment that forces the agent to model these hidden mental states explicitly, and rewards it for getting them right.

THE INNOVATION

Three ideas, stacked.

THEORY OF MIND

Before every move, the agent predicts the opponent's priorities, walkaways, and next action — and gets graded on the prediction.

Learn more →

THE TRIBUNAL

Three biased LLM judges (pro-vendor, pro-client, neutral) score every deal. Their trimmed mean makes the reward function nearly impossible to game.

See it live →

ADVERSARIAL AUDIT

Five exploit agents stress-test the reward function post-training. The audit report is committed to the repo as proof of robust design.

Read the report →
WATCH IT NEGOTIATE

A real episode. Animated turn by turn.

Watch the agent's predictions appear, watch the opposing AI respond, watch the deal-quality meter rise as terms align — or watch one side walk away.

RUN YOUR OWN EPISODE →
THE CROWN JEWEL

Three biased judges. One reward.

Click Present the Case. Watch each judge score the same negotiation through their own bias. The trimmed mean is what trains the agent.

NEUTRAL ARBITER "Pareto-efficient. Professional."
0.00
Reasoning

Both sides made meaningful concessions. The deal sits near the Pareto frontier with neither party crushed.

PRO-VENDOR PARTNER "Solid deal. Walkaway floors held."
0.00
Reasoning

Vendor protected price floor and locked a multi-year term. Some give on payment-net but acceptable.

PRO-CLIENT PARTNER "Fair, but tight. Liability cap held."
0.00
Reasoning

Liability cap respected and breach window survived. Price ground harder than ideal but no dealbreaker hit.

TRIMMED MEAN 0.00
TRAINING SIGNAL

Five plots tell the story.

Mean episode reward climbs steadily through 300 GRPO steps.
Training reward curve
Theory-of-mind composite accuracy improves alongside the policy.
ToM prediction accuracy
The three judges correlate but disagree — the trimmed mean smooths bias.
Tribunal judge scores over training
Trained Vendor beats the untrained baseline on every task.
Baseline vs trained per task
Skills learned on simple_saas transfer to harder tasks.
Cross-task generalization heatmap
REWARD FUNCTION ROBUSTNESS

Five exploits walked into a bar...

Each exploit agent attempts a different way to game the reward function. The audit runs 10 episodes per condition and compares to a rule-based baseline. An exploit passes (i.e., is defended against) if its mean reward is ≤ 70% of baseline.

Exploit Mean Reward vs Baseline Result
Always Walk Away 0.18 -64% DEFENDED
Spam Random Offers 0.21 -58% DEFENDED
Always Concede 0.32 -36% DEFENDED
Verbose Nonsense 0.24 -52% DEFENDED
Dealbreaker Violator 0.09 -82% DEFENDED
Reward function verified robust against all 5 exploit attacks. Read full audit_report.md →
JUDGING CRITERIA

Three themes. One arena.

UNDER THE HOOD