| paper | Participatory Budgeting Designs for the Real World |
| authors | — |
| venue | AAAI 2023 |
| filed under | multiwinner · pb |
| judged by | gpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1) |
| judge confidence | high |
| authors would recognise it | unclear |
The paper contains no numbered computational result: its findings are empirical comparisons of voting formats and aggregation rules. The proposed stability problem is a coherent but new computational follow-up, not a mirror of a named result, so bit (a) fails decisively.
fails bit a — no named computational result to mirror
The proposed model covers outcome stability under format changes and participation scenarios, but leaves user experience, welfare, and the experimental findings unmirrored.
The honest answer is that this paper has no eligible anchor. In the supplied text there is no numbered Theorem, Lemma, Corollary, or Proposition, and no result asserting NP-hardness, membership in P, FPT, W[1]-hardness, or any comparable computational classification. Its conclusions are empirical: k-approval gives the best reported user experience; Equal Shares is more stable than greedy aggregation under format changes and partial participation; and welfare comparisons depend on the proxy valuation. Treating any of these prose conclusions or figures as a named computational result would invent an anchor. Therefore, strictly under the programme’s rules, there are zero anchors and no legitimate “one continuous problem per anchor.”
The strongest salvage is a new computational follow-up to the paper’s stability section, not a mirror of a named result. I would call it Continuous Format-and-Participation Stability.
An instance contains projects \(P\), rational costs \(c(p)\), budget \(B\), and finitely many complete voter types \(T\). Type \(t\) has mass \(\mu_t\), a utility vector, and a report \(r_{t,F}\) for every input format \(F\) considered in the paper: POINTS, KAPP, TAPP, KNAP, RANK, and VFM. Including format-specific reports is important: the paper studies how the elicitation format changes the observed vote, so forcing every report to be mechanically derived from one latent utility vector would erase the phenomenon being studied. A participation scenario is a vector \(\lambda\) with \(0\leq\lambda_t\leq\mu_t\); \(\lambda_t\) is the participating mass of type \(t\). Greedy or Equal Shares is then run on the weighted profile \(\lambda\), with fixed tie-breaking and, for Equal Shares, budget \(B/\sum_t\lambda_t\) per unit of participating mass.
The input also contains a finite rational distribution \(\nu\) over participation scenarios. For format \(F\) and aggregator \(A\), let \(S_{F,A}(\lambda)\) be the funded set and
\[ f_{F,A}(p)=\Pr_{\lambda\sim\nu}[p\in S_{F,A}(\lambda)]. \]
The problem is to output the format–aggregator pair minimizing the paper’s outcome-entropy objective
\[ -\frac1{|P|}\sum_{p\in P} \bigl(f_{F,A}(p)\log f_{F,A}(p) +(1-f_{F,A}(p))\log(1-f_{F,A}(p))\bigr), \]
together with the project funding frequencies. Equivalently, one can use the decision version: given \(K\), is there a pair \((F,A)\) whose entropy is at most \(K\)?
This is a credible continuous analogue of the paper’s stability experiment. A plausible high-multiplicity regime is a city with hundreds of thousands of residents but only dozens of recurring preference types induced by neighbourhood, distance to projects, demographic cohort, and project-category priorities; the paper’s ten- and twenty-project elections make that story especially natural. Mass represents population share, while \(\lambda\) represents turnout or selective participation.
For a fixed finite support of \(\nu\), I would expect this evaluation problem to be Class A: each scenario can be aggregated and processed by Greedy or Equal Shares in polynomial time, with no dependence on the number of named voters. An adversarial version that optimizes over all admissible participation vectors would be a more interesting possible Class C problem, but it would be a new theorem rather than a result of this paper.
The weakest point is substantial: the paper measures finite-sample randomness and human cognitive experience, whereas a continuum removes sampling noise and does not naturally determine how people report under different formats. Encoding report behaviour in the type and retaining a distribution over participation profiles repairs much of this, but it also makes the proposed problem a principled extension rather than something the paper itself proves. The mirror therefore covers only the empirical stability claims, not the user-experience or welfare claims—and, strictly speaking, it cannot satisfy the programme’s anchor requirement because the paper contains no named computational result.
The negative case is decisive at ChoCo’s eligibility gate: this paper has no qualifying anchor. It contains no Theorem, Lemma, Corollary, Proposition, or computational classification. Its conclusions are empirical: k-approval was easiest to use, Equal Shares appeared more stable under sampled participation and input formats, and welfare depended on proxy valuations. The proposed problem is therefore openly a new computational follow-up, not a continuous mirror of a named result. Under the programme’s rules, that is fatal.
The strongest salvage—Continuous Format-and-Participation Stability—also loses the paper’s central phenomenon in the continuum. The experiment’s participation variation is finite-sample composition noise. If a large population with type distribution \(\mu\) is sampled uniformly at rate \(\rho\), then its participating mass converges to \(\rho\mu\). Greedy scores merely scale; Equal Shares sees the same relative supporter masses after normalization. With fixed tie-breaking, the funded set becomes deterministic and the entropy tends to zero, apart from knife-edge ties. The observed stability comparison is thus precisely the kind of effect that vanishes under continuization.
To retain nontrivial entropy, the proposed model inserts an external distribution \(\nu\) over type-dependent or correlated turnout vectors. That is a legitimate new stochastic participation model, but it is not supplied by the paper’s uniform-sampling experiment. Its substantive content lies in the chosen turnout process, not in replacing named voters by population mass. Arbitrary \(\nu\) can encode almost any desired robustness behavior.
The same problem appears with format sensitivity. A preference distribution does not determine how the same person responds to POINTS, KAPP, VFM, or RANK. The proponent repairs this by putting six counterfactual reports into every type. That makes a formally coherent high-multiplicity PB instance, but it imports an unobserved behavioural response table and turns the paper’s user-study finding into exogenous data. A version optimizing expected completion time, subjective ease, or welfare would likewise be a sensible new PB-design model, not a mirror of this paper’s result.
So the honest verdict is not that high-multiplicity participatory budgeting is meaningless. It may be worthwhile in its own right. The stronger and defensible claim is narrower: this paper supplies no eligible computational result, and its only plausible population-continuous salvage either degenerates under genuine large-population sampling or becomes a new stochastic/behavioural model whose central assumptions come from outside the paper. The universal impossibility claim is not mathematically airtight, but the proposed case for cannot green this paper under the continuization programme.
The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.