| paper | Curiosity-Driven Partner Selection Accelerates Convention Emergence in Language Games |
| authors | — |
| venue | AAMAS 2025 |
| filed under | unclassified |
| judged by | gpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1) |
| judge confidence | high |
| authors would recognise it | no |
The paper contains no numbered theorem-level computational result: its contributions are simulation-based comparisons and empirical convergence measurements. Any continuous formulation over learner-state distributions would be a new mean-field or control problem, not a mirror of an eligible result from this paper.
fails bit a — no named computational result to mirror
No eligible result is covered; the empirical partner-selection, network, and convergence findings remain outside the programme's anchoring requirement.
The strongest honest answer is that this paper contains no eligible anchor for a continuous mirror. It has Algorithm 1, experimental claims, figures, and Table 1, but no numbered Theorem, Lemma, Corollary, or Proposition asserting a complexity result such as NP-hardness, membership in P, W[1]-hardness, or FPT. Results cited from [34] and elsewhere are not results proved by this paper.
Therefore there are zero anchors, and no per-anchor continuous problem can honestly be supplied under the programme’s rules. The central empirical claim in Section 4.1—that curiosity-driven partner selection reduces convention-emergence time—could inspire a new mean-field problem over masses of complete learner states, with Boltzmann partner-selection dynamics and an objective of reaching ACC/ACE thresholds quickly. But that would be a proposed research direction, not a mirror anchored in a named result of this paper.
Its main weakness would be type fragmentation: full Q-tables, learning histories, and network neighbourhoods may make almost every agent a distinct type, eliminating the high-multiplicity gain.
The negative case is unusually strong here: this paper supplies no eligible anchor at all. Its contributions are an empirical comparison of two fixed learning procedures, reported through simulations, plots, and convergence-time measurements. There is no theorem, lemma, corollary, or proposition asserting the complexity of a computational problem. The paper neither defines an optimization problem over partner-selection policies nor proves hardness, tractability, approximation guarantees, or parameterized results. Thus there is no named computational result whose high-multiplicity counterpart could be studied.
The obvious rescue is to continuize the claim in Section 4: given a distribution of learner states, study the time until \(95\%\) convention adoption, perhaps comparing curiosity-driven and random matching. But that is precisely a mean-field or stochastic-dynamics programme, not a continuous computational-social-choice mirror. It asks for existence, convergence, rates, or stationary behaviour of a population process. The paper’s own conclusion identifies this as an open “mean dynamics” question. Turning it into a computational problem—say, choosing a policy or hyperparameters to minimize threshold-hitting time—would be a new control problem, not a continuation of a result established here.
Nor does the paper have a useful finite type space hidden beneath its population size. An agent’s relevant state includes its full concept–word \(Q\)-table, its partner-selection \(Q\)-values, and the network neighbourhood through which those values are indexed. The partner-selection table is explicitly about named potential partners. After interactions begin, two agents with the same lexicon can still have different valuations of every possible partner, and those valuations determine future transitions. Histories therefore matter, not merely current lexicons. Replacing agents by masses of lexicon types would discard the mechanism the paper claims is responsible for acceleration.
A more sophisticated mirror could use a distribution over complete learner states together with a measure-valued interaction kernel or graphon, retaining pairwise partner affinities. That avoids the crude objection that “prices belong to individuals”: the complete learner state could indeed define a type. But it also shows why this is not the programme’s high-multiplicity setting. The type is now a continuously valued, function-like object containing a coordinate for a potentially continuum of partners; the interaction kernel and state distribution evolve jointly. One has a measure-valued stochastic process on an infinite-dimensional state space, not a fixed finite set \(T\) with an LP, separation problem, or high-multiplicity algorithmic regime. Allowing the number of distinct complete states to grow with the population merely renames the original individualized simulation.
The network experiments do not repair this. The relevant object is not just a distribution of degrees or graph motifs: the paper’s effect depends on which particular agents are connected, what each remembers about each neighbour, and how those pair-specific values evolve. A graphon limit could be mathematically interesting, but it would be an analytic reformulation of networked reinforcement learning. It would not preserve a computational question from the paper, because the paper never poses one beyond running the prescribed algorithm and measuring its outcome.
There is also no intervention problem to mirror. The authors do not optimize over communication graphs, partner-selection rules, or learning rates in a formal input instance; they select a few graph families and tune parameters experimentally. One could invent an optimization problem over those choices, but that would be a worthwhile new paper perhaps—not a continuous mirror of this one. The proponent’s proposed mean-field direction is therefore best understood as a new research agenda, and it lies outside ChoCo’s stated computational scope.
The honest conclusion is not that a continuum model of language-game learning is impossible. Such a model may be valuable for mean-field analysis or reinforcement-learning theory. It is that this paper contains no theorem-level computational result to continue, and every stronger candidate mirror either loses the identity- and history-dependent partner-selection mechanism or becomes an infinite-dimensional population-dynamics problem unrelated to the programme’s finite-type computational landscape. The proponent’s “zero anchors” assessment therefore defeats the case for a worthwhile ChoCo mirror.
The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.