Learning Partner Selection Rules that Sustain Cooperation in Social Dilemmas with the Option of Opting Out

· AAMAS 2024 (aamas24-00128)

no mirror
paperLearning Partner Selection Rules that Sustain Cooperation in Social Dilemmas with the Option of Opting Out
authors
venueAAMAS 2024
filed underunclassified
judged bygpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1)
judge confidencehigh
authors would recognise ityes

Why no mirror

The paper contains no numbered theorem, lemma, corollary, or proposition asserting a computational complexity or algorithmic result; its claims are simulation findings, so bit (a) fails. A repaired mean-field formulation is plausible, but it cannot supply the missing computational anchor. Therefore the grade is red.

fails bit a — no named computational result to mirror

The objection that survived

The proposed \(16\)-class marginal process is not Markov because persistent pairs create correlations, while faithful learning types must include mutable \(Q\)-tables and matching state.

fatal: False

What the mirror covers

A continuous model could cover the paper's SDOO dynamics, cooperation rates, OFT/TFT emergence, rewards, and forced-OFT comparison, but none is an admissible numbered computational result.

Open questions for a prover

The case FOR (proponent)

The strongest honest positive case is real, but it has no admissible anchor under your rules: this paper contains no numbered Theorem, Lemma, Corollary, or Proposition, and no result asserting membership in \( \mathrm{P} \), NP-hardness, W[1]-hardness, FPT, or any comparable computational classification. Algorithm 1 and equations (1)–(2) are definitions, while the claims in Sections 4.1–4.3 are simulation findings. The Q-learning convergence results are cited from elsewhere and are not computational results proved here.

So there are formally zero anchors. I would not pretend that “cooperation emerges and is sustained” is a theorem.

The best available mirror would nevertheless be a direct mean-field version of the paper’s central SDOO experiment. Call it Mean-Field SDOO Learning Verification.

A type is a complete learner state: its Q-table, its last Prisoner’s Dilemma action, and the fixed learning parameters. Two agents of the same type therefore have identical action probabilities and identical update rules. The society is a distribution \( \mu \) over these types. In a policy-level simplification, the behavioural part of a type is one partner-selection rule
\[ \sigma:\{C,D\}\to\{N,Y\} \]
and one game rule
\[ \rho:\{C,D\}\to\{C,D\}, \]
giving \(4\times4=16\) policy pairs, with the agent’s last action added to the state. Thus the mass of the OFT–TFT type, for example, is the fraction of agents who stay after observing \(C\), switch after observing \(D\), and play TFT in the Prisoner’s Dilemma.

An instance consists of rational Prisoner’s Dilemma payoffs, the horizon \(M\), number of learning episodes \(E\), learning rate \( \alpha \), inverse temperature \( \kappa \), discount factor \( \gamma \), an explicitly represented initial distribution \( \mu^0 \), and a cooperation threshold \( \eta \). The continuum process applies Algorithm 1 to \( \mu \): pairs are formed according to the product distribution, pairs dissolve when either member chooses \(Y\), dissolved mass is rematched according to its conditional distribution, and the two agents then play and update their Q-tables exactly as in equation (1). This defines a deterministic mean-field evolution
\[ \mu^{r+1}=\Phi(\mu^r). \]

The problem is to compute an \( \varepsilon \)-accurate final distribution and decide whether the expected fraction of cooperative outcomes, or the expected average population reward, after \(E\) episodes is at least \( \eta \). A solution is the returned trajectory or final distribution together with an \( \varepsilon \)-accurate value of the chosen aggregate statistic.

This is a credible mirror because the paper already studies population percentages, aggregate reward, random matching, and distributions over the \(4\times4\) strategy classes. A natural regime is a very large population of cloned or cohort-based software agents—same payoff matrix, observation model, learning algorithm, and hyperparameters—repeatedly paired by a platform. Then \(N\) is enormous while the number of behavioural types is small, or at least much smaller than \(N\). Continuity is applied to the society, not to the payoff outcomes or to the randomisation mechanism.

For the fixed-policy, finite-horizon version, I would expect a Class A result: the pairwise transition map can be evaluated by dynamic programming over type pairs, with cost polynomial in the number of types and the horizon. The fully faithful Q-learning version is less clear: Q-tables become part of the type, and their support can grow rapidly with the number of episodes. That version may be continuum-specifically hard, or may require approximation. Further questions include convergence to a stationary distribution, robustness under trembling-hand or \( \varepsilon \)-greedy exploration, finite-\(N\) error bounds, and the minimum mass of OFT or TFT learners needed to sustain a cooperation threshold.

The weakest point is that replacing the learner state by the paper’s \(16\) behavioural classes may no longer be the authors’ learning problem: Q-values, not merely strategy labels, determine future behaviour. Including Q-values restores fidelity but weakens the small-type high-multiplicity story. Moreover, the paper supplies no law-of-large-numbers argument showing that its finite simulations converge to this mean-field process. Thus this is a promising continuous research problem inspired directly by the paper, not a continuous analogue anchored by one of its named computational results.

The case AGAINST (opponent, writing after the proponent)

The negative case is decisive at the programme’s first gate: this paper supplies no admissible computational anchor. It proves no numbered theorem, lemma, corollary, or proposition about the complexity or algorithms of a computational problem. Algorithm 1 and equations (1)–(2) define the simulation; the claims about cooperation are empirical observations. The Q-learning convergence results are imported from cited work, and their hypotheses are not established for this co-learning environment. Thus there is no theorem here whose high-multiplicity analogue could be classified.

The proposed Mean-Field SDOO Learning Verification does not repair that absence. As stated, its \(16\) policy classes do not define a Markovian population state. In SDOOs, pairs that both choose to stay remain together. Consequently, the next transition depends on the joint state of a pair—both learners’ last actions, \(Q\)-tables, and current policies—not merely on the marginal distribution \(\mu\). After the first round, the pair distribution is not the product of marginals. A faithful continuum state would have to be a measure \(\nu\) over matched learner states, with rematching determined by the mass released from each pair.

Nor are the \(16\) strategy labels sufficient for the learning process. Under Boltzmann exploration, two agents classified as OFT or TFT can have different action probabilities and different future updates because their \(Q\)-values differ. Exact types must therefore include the mutable \(Q\)-table, learning parameters, last action, and matching state. Identical software agents rapidly separate into many such states through stochastic experience. Collapsing them to \(4\times4\) labels preserves the paper’s plots but removes the mechanism the paper claims to discover. Retaining the full state produces a history- and episode-dependent type space whose support can grow exponentially, rather than the fixed finite-type high-multiplicity object that makes continuization computationally useful.

The obvious repairs do not yield a mirror of this paper. A fixed-policy model over pair states is a legitimate mean-field dynamical system, but it deletes Q-learning and studies a different object. A full learner-state model is faithful, but is a generic measure-valued analysis of multi-agent learning. Asking for the minimum initial mass of OFT/TFT agents, or the smallest perturbation that destroys cooperation, could create an interesting mean-field control or robustness problem; it would introduce a new intervention problem absent from the paper, rather than continuize one of its computational results.

This is not the objection that continuity “does not help.” The problem is prior: the paper never formulates the relevant computational question, and its central phenomenon cannot be represented by the proposed small type space without discarding the phenomenon itself. The cloned-software-agent scenario does make a continuum population scientifically conceivable, so the universal negative is not airtight as a claim about all future research. But under ChoCo’s stated standard, the strongest honest verdict is that this paper offers no worthwhile continuous computational mirror—only a possible mean-field-learning research agenda that is adjacent to, but not an instance of, the programme.

The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.