| paper | Joint Scoring Rules: Competition Between Agents Avoids Performative Prediction |
| authors | Rubi Hudson |
| venue | AAAI 2025 |
| filed under | unclassified |
| judged by | gpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1) |
| judge confidence | medium |
| authors would recognise it | yes |
Theorem 6
statement extracted from the paper’s text layer
Given \(A\), \(O\), finite \(T\), rational \(\mu\in\Delta(T)\), a strict principal preference, a symmetric strictly proper scoring rule \(s\), and hidden conditional distributions \(q_a\), each query elicits type-level conditional reports for a binary action partition through mass-weighted LSS scoring and \(D_\mu^{\mathrm{ISA}}\); output the unique \(a^\star\) satisfying \(q_{a^\star}\succ q_a\) for all \(a\ne a^\star\), while minimizing worst-case queries and determining whether \(O(\log |A|)\) comparisons suffice uniformly over \(\mu\).
A high-multiplicity society of predictive-agent types, with type masses \(\mu_t\), conditional forecasts \(p_{t,a}\in\Delta(O)\), mass-weighted zero-sum LSS scores, \(D_\mu^{\mathrm{ISA}}\), and adaptive branch comparisons whose objective is to identify \(a^\star\).
The \(O(\log |A|)\) bound is insensitive to the numerical masses, and treating a positive-mass type as a strategic bloc changes the deviation game; it remains open whether the paper's equilibrium reasoning supports this exact continuum mechanism.
fatal: False
The mirror covers Theorem 6's binary-search identification guarantee and leaves the incentive, uniqueness, stochastic-choice, and experimental results outside its scope.
There is a credible, deliberately narrow positive case. I would anchor it on Theorem 6, rather than claim that the paper’s entire incentive theory already has a continuous analogue.
Consider a large deployment of predictive agents advising a hospital, regulator, or AI system before it chooses one intervention. There are \(N\) deployed agents but only \(\tau\) complete types, with \(\tau\ll N\). A type includes the model architecture, training data, information access, calibration convention, and scoring parameters; agents of the same type are indistinguishable for the mechanism. The population is represented by \(\mu\in\Delta(T)\), where \(\mu_t=n_t/N\).
The action set is \(A\), the outcome set is \(O\), and \(q_a\in\Delta(O)\) is the true outcome distribution conditional on action \(a\). The principal has a strict preference over outcome distributions and wants
\[ a^\star\in\arg\max_{a\in A} q_a. \]
A type-\(t\) predictor reports \(p_{t,a}\in\Delta(O)\) conditional on every action. The continuum version of the paper’s zero-sum competition can use
\[ S_t(a,p,q) = s(p_{t,a},q_a) - \sum_{u\in T}\mu_u s(p_{u,a},q_a), \]
where \(s\) is a symmetric strictly proper scoring rule. The population-weighted scores sum to zero. In a finite society this is the limit of the paper’s finite-agent LSS rule as repeated types become highly numerous.
For an exact analogue of the paper’s ISA-max mechanism, define
\[ D_\mu^{\mathrm{ISA}}(p)=a \]
when some positive-mass type reports a distribution for \(a\) that is at least as preferred by the principal as every report for every other action, with the paper’s tie-breaking rule applied. A genuinely mass-sensitive companion would instead use
\[ \bar p_a=\sum_{t\in T}\mu_t p_{t,a} \]
and choose the action maximizing \(\bar p_a\). I would treat that mean-max version as a new extension, not pretend that the paper proves it for arbitrary \(\tau\).
The anchor is Theorem 6, proved in this paper:
“A principal can identify \(a^\star\) with at most \(O(\log(|A|))\) comparisons between actions.”
My continuous problem would be called Continuum Conditional-Action Search. Its instance consists of \(A\), \(O\), a finite type set \(T\), rational masses \(\mu_t\), the principal’s preference, the scoring rule \(s\), and the same conditional-comparison interface used by Algorithm 1: for a current split of the candidate actions, each type supplies predictions conditional on the two branches. The task is to output the exact action \(a^\star\), together with a valid comparison transcript, while using as few branch comparisons as possible. A truthful solution must operate on the type-level report field and remain valid under the intended LSS equilibrium.
This is a legitimate high-multiplicity mirror. The original paper’s objects remain intact: agents make conditional predictions, the principal selects an action, predictions can manipulate that selection, and zero-sum competition is the proposed remedy. Only the population representation changes from \(N\) named predictors to the distribution \(\mu\) over recurring predictor types. The scenario is also plausible: large forecasting ensembles, replicated model deployments, or many copies of a small number of model/data-source types. It is not a claim that every collection of predictors has high multiplicity; this is a specific deployment regime.
I expect the explicit-type version to be Class A. Given truthful type reports, evaluating the continuum mechanism requires only weighted sums over \(\tau\) types and \(|O|\) outcomes, hence polynomial time in \(\tau\), \(|O|\), and the input bit length. Under the paper’s comparison interface, Theorem 6 suggests an \(O(\log|A|)\)-query implementation. If \(A\) is given only implicitly, the interesting question becomes whether the branch-prediction oracle itself has an efficient separation algorithm; that could produce a continuum-specific hard variant.
I use only this anchor. Theorem 7, also proved here, gives \(O(1)\) comparisons, but it relies on an additional nonzero-sum pilot prediction and is best viewed as a refinement of the same search result, not an independent population phenomenon. Theorems 2–5 and 8–10 are useful support, especially the LSS and mean-max results, but they are mechanism and incentive theorems rather than complexity classifications.
The main further questions are whether the LSS equilibrium proof survives nonatomic or type-coalition deviations; whether the mass-weighted mean-max rule remains truthful for arbitrary \(\mu\); how heterogeneous information \(q_t\) should be scored; and whether noisy or approximate reports preserve the search bound.
The weakest point is real: under ISA-max, the decision rule depends mainly on the support of \(\mu\), not on the precise masses. Moreover, an infinitesimal individual cannot usually pivot the decision, so the strategic game is not automatically the same as the paper’s finite-\(n\) game. A referee could therefore regard this as a high-multiplicity computational mirror with a new equilibrium problem, rather than a theorem transfer. The case nevertheless survives because the population regime is natural, the conditional-prediction problem is unchanged, and Theorem 6 supplies a concrete algorithmic question whose continuous representation is meaningful and plausibly tractable.
The proponent’s sole anchor, Theorem 6, does not survive as a worthwhile ChoCo mirror. The theorem is a query bound for identifying \(a^\star\); it is not a computational result whose difficulty depends on the population. Algorithm 1 works for two named predictors, for \(N\) predictors, and for any repeated-type expansion. Replacing the \(N\) agents by \(\mu\) therefore adds no computational object for the population to determine.
This is especially clear for the proposed \(D_\mu^{\mathrm{ISA}}\). It depends on \(\mu\) only through its support:
\[ \operatorname{supp}(\mu)=\{t:\mu_t>0\}. \]
Two societies with the same supported types but radically different masses induce the same decision rule and the same \(O(\log |A|)\) search. The precise population distribution is dead data. The score expression contains \(\mu\), but Theorem 6 does not optimize, search over, or otherwise compute with those masses; it merely repeats the paper’s existing comparison protocol.
The continuum also damages the strategic content that makes the paper interesting. In an atomless population, one predictor’s deviation has measure zero, so it cannot manipulate a mass-based mean or aggregate. The performativity problem disappears before zero-sum competition is used. Truthful reporting may remain optimal under a proper scoring rule, but only because the agent cannot affect the chosen action. Under ISA-max, the situation is worse: a type with any positive mass can matter while its exact mass is irrelevant.
The strongest mass-sensitive repair is weighted mean-max,
\[ \bar p_a=\sum_t \mu_t p_{t,a}. \]
But this turns the population into one pooled forecast. Theorem 6 then becomes an ordinary binary-search result supplied with an aggregation subroutine. If positive-mass types are instead allowed to coordinate their reports, the model becomes a coalition game. That is not the finite-agent equilibrium in the paper: cloning a type creates many individually deviating agents, not one strategic bloc. The paper’s own Theorem 3 is restricted to \(n=2\), and it explicitly warns that the mean-max argument fails for general \(n\). A coalition-proof continuum extension could be interesting, but its interest would come from a new mechanism-design problem, not from Theorem 6.
The proposed forecasting-ensemble scenario does not repair this. If model architecture, data, information, and calibration are all part of a complete type, genuinely different predictors become different types; if they are not, the aggregation forgets precisely the information that makes them distinct. In the first case \(\tau\) approaches the number of named agents; in the second, the population collapses to repeated copies whose multiplicity has no role in the theorem.
Theorem 7 does not help: its \(O(1)\) search is the same action-query result with an additional pilot prediction. Thus the faithful mirror is merely a reparameterized implementation theorem, while the mass-sensitive mirrors are new coalition or forecast-aggregation models for which this paper supplies no computational result. The negative case is not a proof that no future continuum mechanism could be devised, but it defeats the only proposed anchor as a genuine continuous-population problem.
The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.