Local Justice and Machine Learning: Modeling and Inferring

· AAAI 2023 (aaai23-25737)

no mirror
paperLocal Justice and Machine Learning: Modeling and Inferring
authors
venueAAAI 2023
filed underunclassified
judged bygpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1)
judge confidencehigh
authors would recognise itno

Why no mirror

Bit (a) fails: the paper has no numbered theorem, lemma, corollary, or proposition asserting a computational complexity or algorithmic result. The proposed finite-type LP is coherent, but it introduces aggregate moral stakeholders and replaces learning one stakeholder’s reward with forward policy optimization. Thus it is a worthwhile new model, not a mirror of a result in this paper.

fails bit a — no named computational result to mirror

The objection that survived

The proposed problem adds a distribution over moral stakeholder types and changes the task from inferring an individual reward function to optimizing known aggregate moral reward; patient mass merely parametrizes newly specified transitions.

fatal: True

What the mirror covers

The proposal covers only a new forward allocation extension and leaves individual reward inference, active query selection, Bradley–Terry learning, simulations, and experiments without an eligible computational result.

Open questions for a prover

The case FOR (proponent)

The strongest honest conclusion is that this paper has no eligible anchor. It contains no numbered Theorem, Lemma, Corollary, or Proposition, and no named computational claim of NP-hardness, membership in P, FPT, W[1]-hardness, or approximation. Its contributions are a mathematical model, a Bayesian active-learning procedure, simulations, and a small human-subject study. Thus, under the programme’s rules, there is no paper result to which a continuous computational mirror can be anchored.

There is nevertheless a credible candidate mirror, though I would label it explicitly as an extension rather than a result of the paper.

The lead candidate is Finite-Type Continuum Dynamic Moral Allocation. Let \(Z\) be a finite set of complete demographic/health types, each encoding all features relevant to the epidemic and allocation process. The society is a rational distribution \(\mu\in\Delta(Z)\); \(\mu_z\) is the mass of type \(z\). A state records cured, susceptible, and deceased mass for each type. At each phase \(t\), the planner chooses a continuous allocation \(a_{t,z}\), subject to capacity and availability constraints. The deterministic transition is a rational affine map \(s_{t+1}=F_t(s_t,a_t)\), representing the large-population approximation used by the paper.

Stakeholders are also divided into finitely many complete moral types \(\Theta\). A moral type \(\theta\) has weights \(w_{\theta i}\) and saturation thresholds \(c_{\theta i}\), with reward

\[ R_\theta(s)=\sum_i w_{\theta i}\min\{x_i(s),c_{\theta i}\}, \]

exactly reflecting the paper’s equation (1). Given a rational distribution \(\nu\) over stakeholder types, the problem is:

Given \((Z,\mu,\Theta,\nu,s_0,F_1,\ldots,F_H)\), find a feasible allocation sequence \(a_1,\ldots,a_H\) maximizing the aggregate discounted moral reward
\[ > \sum_{\theta\in\Theta}\nu_\theta > \sum_{t=1}^H\gamma^t > \bigl(R_\theta(s_{t+1})-R_\theta(s_t)\bigr). > \]

A solution is the allocation sequence together with its exact objective value, or an \(\varepsilon\)-optimal sequence. For \(0<\gamma\le1\), affine transitions, and nonnegative weights, the natural finite-horizon version is an LP after introducing epigraph variables for the \(\min\) terms. It therefore plausibly belongs to Class A. Its input size depends on the number of demographic and moral types, not on the number of individual residents.

The regime is plausible in national or regional public-health planning: millions of residents, but relatively few complete administrative feature bundles, and a small number of recurring stakeholder moral profiles used by hospitals, patient groups, unions, and public-health organizations. This is not merely declaring every individual interchangeable; it specifies the attributes relevant to the model and groups agents sharing all of them. The paper itself supports the regime: it models allocations as proportions and explicitly motivates deterministic transitions by the population being large.

The authors might recognize this as a continuous extension of their allocation model, especially because they discuss collective responses and aggregate rewards. But the weakness is substantial: their actual learning problem concerns one stakeholder’s unknown \((w,c)\), whereas the proposed problem aggregates a distribution \(\nu\) of stakeholders and optimizes policy. A referee could reasonably call that a new social-choice problem rather than the paper’s own computational problem. Moreover, the paper gives no formal transition complexity or optimization guarantee.

Further questions would include continuum preference inference from aggregate Bradley–Terry responses, robust allocation when \(\nu\) is uncertain, and minimax or fairness-constrained objectives across moral types. These could yield genuine Class A, B, or C questions. But they cannot repair the central evidentiary gap: this paper supplies a promising model to continuize, not a named computational result that the programme can mirror.

The case AGAINST (opponent, writing after the proponent)

On the programme’s own eligibility rule, this paper fails before the modelling question begins. It contains no numbered theorem, lemma, corollary, or proposition asserting the complexity of a computational problem. Its contributions are a preference model, a Bayesian elicitation heuristic, simulations, and a 33-person study. The MDP is a formal representation, not a computational result; the query-selection procedure is described as a heuristic, not analysed for complexity or approximation. Consequently there is no named result to mirror.

The proposed “Finite-Type Continuum Dynamic Moral Allocation” does not repair that defect. It changes the paper’s problem in two decisive ways.

First, it continuizes the wrong population. The people being allocated cures are state variables, not preference-bearing agents. Their masses already appear in the paper’s state and action descriptions. The agents whose preferences matter are stakeholders, and the paper models one stakeholder’s unknown \((w,c)\) at a time. Introducing a distribution \(\nu\) over moral types therefore adds a new population and a new aggregation problem; it is not a high-multiplicity version of the paper’s inference task.

Second, the proposed objective reverses the direction of the paper. The paper asks the planner to infer an individual stakeholder’s reward function from pairwise trajectory choices. The proposed LP instead assumes a distribution of known moral types and chooses an allocation maximizing their aggregate reward. That is a forward social-planning problem, not preference elicitation. The paper explicitly leaves aggregation of preferences outside its scope. Calling this an extension is fair; calling it a continuous mirror is not.

The public-health setting does make a large-population regime plausible. Millions of patients can reasonably be represented by masses of complete demographic and health types. But that only validates the modelling regime; it does not provide the missing computational anchor. In the paper, population size has already been eliminated by the deterministic proportional state description. There is no individual-level allocation problem, no population-size parameter whose high-multiplicity relaxation can be compared with a discrete version, and no discrete computational theorem from which a continuous counterpart could inherit meaning.

A more faithful mirror would retain one stakeholder and replace the synthetic patient population by a distribution over complete feature types. That is a legitimate model, but the population distribution merely determines the transition map and initial state. Once those affine transitions are supplied, the allocation problem is an ordinary finite-horizon optimization problem. The computational content comes from the newly specified transition model, not from continuization. If the transitions are not supplied, one must newly define how overlapping groups, correlations, infection, and allocation interact; that is substantial new modelling, not a mirror of the paper.

The same problem appears on the learning side. Pooling many respondents into a distribution of moral types could motivate mixture inference or population-level Bayesian experimental design. But the paper’s object is the preference of each respondent. Aggregating responses loses the individual parameter that the method is designed to infer. Grouping respondents with identical \((w,c)\) is not intrinsically invalid—the programme is right that high multiplicity concerns complete types—but the paper neither posits such repeated moral types nor studies their aggregate decision problem. Adding them creates a new computational social-choice paper.

Thus the negative case is very strong as a programme-screening verdict: this paper offers no eligible computational result, and the strongest proposed mirror is a coherent but unanchored forward-planning extension. The universal claim that no imaginable adjacent scenario could ever be worthwhile is weaker; one could certainly write a sensible paper about finite-type public-health allocation. But that would be a new ChoCo problem inspired by this paper, not a continuous mirror of a result it contains.

The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.