TDG4Crowd:Test Data Generation for Evaluation of Aggregation Algorithms in Crowdsourcing

Yili Fang, Chaojie Shen, Huamao Gu, Tao Han, Xinyi Ding · IJCAI 2023 (ijcai23-00333)

no mirror
paperTDG4Crowd:Test Data Generation for Evaluation of Aggregation Algorithms in Crowdsourcing
authorsYili Fang, Chaojie Shen, Huamao Gu, Tao Han, Xinyi Ding
venueIJCAI 2023
filed underfrontier · tools-data
judged bygpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1)
judge confidencehigh
authors would recognise itno

Why no mirror

The paper contains no numbered theorem, lemma, corollary, or proposition asserting a computational result; Definition 1 and empirical VAE experiments do not satisfy bit (a). The proposed Population-DDG formulation adds essential modelling choices and therefore cannot supply a qualifying mirror of a result in this paper.

fails bit a — no named computational result to mirror

The objection that survived

No qualifying anchor exists; the strongest objection is that Population-DDG is an added optimization model whose generator family, novelty constraints, and population types are absent from the paper.

fatal: True

What the mirror covers

No named computational result is covered; the candidate extension concerns the paper's distribution-oriented generation objective and leaves its empirical VAE comparisons outside the programme's scope.

The case FOR (proponent)

The strongest honest conclusion is that this paper has no qualifying anchor.

Its only numbered formal statement is Definition 1, “Distribution-oriented Data Generation problem—DDG.” It defines

\[ g^\ast=\arg\max_g \sigma(Q,g(Q)), \]

but gives no theorem, lemma, corollary, or proposition asserting a complexity or algorithmic result. The claims that TDG4Crowd achieves lower KL/KS divergence and more consistent aggregation rankings are empirical experiments, not named computational results. Results cited from other papers cannot serve as anchors for this paper.

Therefore, under the programme’s rules, there are zero anchors and no honest Class A/B/C verdict for a mirror of this paper. One should not manufacture a “Theorem 1” from Figure 2 or treat the VAE training procedure as a proved polynomial-time result.

The best non-anchored candidate would extend Definition 1 as follows. Let \(X\) be finitely many complete task types, \(A\) finitely many complete annotator types, and \(Y\) the label set. A high-multiplicity crowdsourcing population has task mass \(\mu\in\Delta(X)\), annotator mass \(\nu\in\Delta(A)\), assignment mass \(\rho\in\Delta(X\times A)\), and a per-type annotation channel \(h_{x,a,y}\). Given a rational seed distribution \(\widehat p\) over \(X\times A\times Y\), balance constraints, and an admissible generator family \(\mathcal H\), the continuous problem is:

\[ \textsc{Population-DDG}: \quad \min_{(\mu,\nu,\rho,h)\in\mathcal H} D_{\mathrm{KL}}(\widehat p\|q), \qquad q_{x,a,y}=\rho_{x,a}h_{x,a,y}. \]

The solution is the synthetic population law itself, not a finite sample. This is recognisably the paper’s question: many near-identical task and worker cohorts replace individually listed tasks and annotators, while \(\rho\) preserves redundancy and \(h\) preserves noisy annotations. With a convex, efficiently separable \(\mathcal H\), it would plausibly be Class A; with the paper’s unrestricted neural VAE parameterization, nonconvex training may instead be continuum-specifically hard.

That candidate is sensible, but it is supported only by Definition 1—not by a named computational result. Its weakness is precisely that the paper never establishes the computational problem the continuization programme requires.

The case AGAINST (opponent, writing after the proponent)

The negative case is unusually strong because the paper supplies no qualifying anchor at all. Definition 1 is merely

\[ g^\ast=\arg\max_g \sigma(Q,g(Q)), \]

with no fixed computational model, complexity claim, approximation guarantee, or named algorithmic theorem. The VAE procedure and the KL/KS comparisons are empirical claims. Results cited from Dawid–Skene, GLAD, IRT, and others belong to different papers. Under ChoCo’s rules, there is therefore no computational result here whose continuous mirror can be assessed.

The proposed Population-DDG formulation does not rescue the paper. If the generator family contains the empirical distribution \(\widehat p\), then choosing \(q=\widehat p\) gives \(D_{\mathrm{KL}}(\widehat p\|q)=0\). The continuous solution simply reproduces the observed seed distribution; it generates no genuinely new tasks, annotators, or annotations. Preventing this requires novelty, support, distance, or generalization constraints that the paper never defines. Once those are added, the computational problem depends entirely on the chosen family \(\mathcal H\): one choice may yield convex optimization, another nonconvex learning, and another an NP-hard fitting problem. That would be a new model, not a continuization of a result in this paper.

There is also a substantive multiplicity problem. A complete task type for LabelMe or Relation must include the task’s content or feature representation, because that content is precisely what affects difficulty and annotation behaviour. Images and sentences are generally one-off objects, so complete types have essentially unit multiplicity. Coarsening them into latent VAE classes makes the masses meaningful, but then the types are no longer complete: distinct tasks with different annotation channels have been merged. The same problem applies to annotators if their individual response histories matter to the aggregation algorithm.

The assignment mass \(\rho\) does not fix this. It records aggregate task–annotator interaction, but not the finite bipartite incidence structure: degrees, overlaps, isolated vertices, or which workers jointly label which tasks. Those details affect DS, HDS, GLAD, and majority voting, and are central to the paper’s sparsity and redundancy claims. Preserving them requires types containing increasingly large neighbourhoods or graph motifs; in real datasets those types again become nearly unique. A graphon or population-confusion-matrix model could be interesting, but it would be a new statistical theory of crowdsourcing rather than the continuous mirror of TDG4Crowd.

Thus the appended candidate either collapses to exact reproduction, or introduces the missing modelling choices that determine the entire problem. The broader claim that no valuable continuous crowdsourcing model could ever be designed is not provable—one could build a worthwhile graphon or cohort-based programme—but this paper gives no computational result and no well-posed high-multiplicity object from which such a programme can start.

The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.