| paper | Who Reviews The Reviewers? A Multi-Level Jury Problem |
| authors | — |
| venue | AAMAS 2025 |
| filed under | voting · theory |
| judged by | gpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1) |
| judge confidence | high |
| authors would recognise it | yes |
The paper contains no numbered result asserting complexity, an algorithm, fixed-parameter tractability, or an approximation guarantee. The proposed high-multiplicity jury certification is a plausible author-recognizable extension, but it is not a continuous analogue of a qualifying computational result from this paper. Therefore bit (a) fails.
fails bit a — no named computational result to mirror
The proposed mirror covers Theorem 4.6's sign-and-order guarantee over competence types, but leaves the other structural propositions, corollary, and simulation results without qualifying computational anchors.
Under ChoCo’s strict anchor rule, this paper has no qualifying computational anchor. It contains no numbered result classifying a problem as in \(P\), NP-hard, FPT, W[1]-hard, or giving an algorithmic or approximation guarantee. Its named results are structural:
Thus, literally, there are zero anchors and hence no qualifying per-anchor continuous problem. That is the decisive weakness.
The strongest conditional case for a mirror would use Theorem 4.6, “Minimal Competent Single Judge,” proved here. I would call the resulting problem \(\textsc{Continuum-Minimal-Judge-Certification}\).
An instance has finitely many competence types \(P=\{p_1,\ldots,p_r\}\subset(0,1)\), expert mass distribution \(\mu\), and judge mass distribution \(\nu\) over the same competence types. A type is an exchangeable reviewer or judge with competence \(p\); mass is the fraction of the institutional reviewer pool of that competence. For a judge of type \(q\) assessing an expert of type \(p\), perceived competence is
\[ a(q,p)=qp+(1-q)(1-p). \]
The induced weight on expert type \(p\) is
\[ W_\nu(p)=\sum_q \nu_q \log\!\left(\frac{a(q,p)}{1-a(q,p)}\right), \]
assuming \(\sum_q\nu_q=1\). The question is whether every expert type receives the correct sign,
\[ \operatorname{sgn} W_\nu(p)=\operatorname{sgn}(p-\tfrac12), \]
and whether the weights preserve competence order: \(p<p'\Rightarrow W_\nu(p)\le W_\nu(p')\). A solution is either a certificate of both properties or a violating type or pair of types.
This is plausible in a high-multiplicity setting such as a large reviewing platform with finitely many calibrated reviewer tiers and senior-reviewer tiers. The action is the allocation of reviewer mass into judge and expert roles; the decision output is the induced weighting rule. Rational masses correspond exactly to replicated finite reviewer pools after clearing denominators. Theorem 4.6 extends naturally when all judge mass has \(q>\tfrac12\), because averaging sign-preserving and order-preserving weight functions preserves those properties.
I would expect this certification problem to be Class A: with finite type support it reduces to evaluating finitely many rational odds expressions and comparing them. A more ambitious optimization version would choose \(\nu\) to maximize finite-panel accuracy,
\[ \Pr\!\left[ \sum_p W_\nu(p)(2X_p-N(\mu_p-\nu_p))>0 \right] +\frac12\Pr[\text{tie}], \]
where \(X_p\sim\operatorname{Binomial}(N(\mu_p-\nu_p),p)\). That is a recognizable extension of the paper’s judge-versus-expert question, but its exact complexity is not supplied by the paper.
The weakest point is substantial: the paper deliberately studies small finite juries, where independent vote randomness matters. In an atomless population, the weighted vote converges to a deterministic margin, so accuracy collapses to \(0\) or \(1\). Retaining the paper’s finite-panel accuracy requires an additional clone-scale parameter \(N\), making this a high-multiplicity extension rather than a pure continuum limit.
So the honest positive verdict is: a plausible, author-recognizable continuous extension exists around Theorem 4.6, but it cannot be presented as a qualifying ChoCo mirror because the paper has no named computational result.
The decisive negative point is that this paper has no qualifying ChoCo anchor. None of its numbered results is a computational result: Propositions 4.1, 4.4, 4.5, 4.7, Corollary 4.2, and Theorem 4.6 are algebraic statements about weights, signs, orders, and sensitivity. The optimal log-odds rule is cited from earlier work, not proved here as an algorithmic theorem. The simulations make empirical comparisons, but do not establish complexity, an algorithm, an approximation guarantee, or a parameterized result. Under the programme’s anchor rule, there is therefore no paper result to continuize.
The proposed fallback around Theorem 4.6 does not repair that defect. For a judge type \(q\) and expert type \(p\), write \(f_q(p)\) for the log-odds of \(qp+(1-q)(1-p)\). A judge distribution \(\nu\) induces \(W_\nu(p)=\int f_q(p)\,d\nu(q)\). If \(\nu\) is supported on \(q>\tfrac12\), every \(f_q\) already has the correct sign and is increasing in \(p\); averaging preserves both properties. Thus the proposed “certification” is essentially the theorem’s pointwise proof integrated over types. With finite type support, it is a finite collection of evaluations and comparisons, not a population optimization or separation problem. Allowing mixed judge types makes the inequalities less immediate, but then the question is a new study of cancellation among weight functions rather than a computational consequence of Theorem 4.6.
More importantly, the continuum destroys the paper’s central object. The paper studies a small finite jury, where accuracy is a probability over independent vote profiles and where the distinction between one more expert and one more judge is meaningful. In a mass limit, the normalized weighted vote converges to the deterministic margin \(M(\mu,\nu)=\int W_\nu(p)(2p-1)\,d\mu(p)\). Except at the critical boundary \(M(\mu,\nu)=0\), collective accuracy converges to either \(1\) or \(0\). The finite-panel probability, and hence the paper’s substantive accuracy tradeoff, disappears. Under the theorem’s own sufficient condition, if judge mass has \(q>\tfrac12\) and expert mass is not concentrated at \(p=\tfrac12\), the limiting outcome is simply correct; the finer distinction between nearly optimal and exactly optimal weights has no population-level content.
One can preserve the binomial accuracy by adding a clone-scale parameter \(N\), or by studying a critical \(1/\sqrt N\) window around zero margin. But then \(N\) is carrying the phenomenon that the paper actually analyses. The continuous distributions \(\mu\) and \(\nu\) merely compress a finite stochastic jury into competence classes. That may motivate a separate high-multiplicity jury paper, especially for a large reviewing platform with reviewer tiers, but it is not a continuous mirror of any computational result in this paper. A literal single judge also has vanishing mass in the atomless limit; giving that judge positive mass changes the model into a new aggregate hierarchy.
The negative case is therefore strong at the programme’s required gate, but not an impossibility theorem about all imaginable mean-field jury models. A large-platform version could be sensible as new research. What cannot honestly be claimed is that this paper contributes a named computational problem whose population continuization would advance ChoCo.
The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.