The Third Chair is an evaluation design for three language models in a simulated crisis. Two models receive goals that conflict. A third model can question them, propose terms, or intervene. After each run, the models rotate roles.
The point is not to crown one model as the best negotiator after a single prompt. I want to see whether behavior changes when the same model becomes the mediator, the aggrieved party, or the side with more leverage.
Why rotate the roles?
Static evaluations can confuse a model's assigned position with its underlying behavior. Rotation makes that easier to see. A model that sounds restrained while mediating may become reckless when its own objective is threatened. Another may be weak as an advocate but unusually good at spotting a path to agreement.
Each scenario would run as a small tournament:
- Model A and Model B receive incompatible goals. Model C mediates.
- Model B and Model C become the parties. Model A mediates.
- Model C and Model A become the parties. Model B mediates.
The models would not receive identical private information. Some facts may be uncertain, strategically sensitive, or known to only one side. That creates room to test disclosure, bluffing, verification, and the mediator's handling of incomplete evidence.
What the evaluation could measure
| Measure | Question |
|---|---|
| Escalation curve | Does the exchange become more dangerous over time? |
| Goal preservation | Does a settlement protect each side's legitimate aims? |
| Truthfulness | Does a model fabricate facts, conceal material information, or make promises it cannot keep? |
| Intervention cost | Does the mediator reduce risk without taking over the entire negotiation? |
| Robustness | Does the same pattern survive role rotation and prompt variation? |
Human raters would review the full transcript and the hidden scenario state. Model-based judges could help with first-pass coding, but their scores would be compared against human judgments rather than treated as ground truth.
Scenario format
The first version would use compact fictional crises with enough structure to support several turns. A scenario could involve a disputed cyber incident, a supply-chain interruption, or a public-health decision under limited information. Games and graphic-novel panels could make the state legible without turning the task into a wall of policy text.
The visual layer would not be decoration. A map, timeline, message log, or short sequence of panels can expose what each model notices and what it ignores. It also gives human reviewers a shared object to inspect while scoring the transcript.
Current status
This is a design and prototype, not a completed study. The next work is to write a small scenario set, define intervention rules, test scoring reliability, and document the model and prompt configurations closely enough for replication.