Catfished! Impacts of Strategic Misrepresentation in Online Dating

· AAMAS 2024 (aamas24-00117)

no mirror
paperCatfished! Impacts of Strategic Misrepresentation in Online Dating
authors
venueAAMAS 2024
filed underunclassified
judged bygpt-5.6-luna / xhigh (triple__luna__xhigh__c2r1)
judge confidencehigh
authors would recognise itunclear

Why no mirror

The paper contains no numbered computational result: its equations are model definitions and its findings are empirical simulation outcomes. The proposed continuous dating-evaluation problem is a prospective extension, not a mirror of a named result in the paper. The opponent therefore defeats the only route to a qualifying grade.

fails bit a — no named computational result to mirror

The objection that survived

The paper's queues, no-repeat histories, and \(O(1)\) per-user interaction limits do not aggregate into a population distribution without adding a new mean-field kernel or changing the scaling regime.

fatal: True

What the mirror covers

The proposed mirror covers population-level welfare comparison under truthful and deceptive reports, but no named computational result from the paper; it leaves the simulation framework's other strategies, matchers, and stochastic dynamics unanchored.

Open questions for a prover

The case FOR (proponent)

The strongest honest case is limited: this paper has no qualifying named computational anchor. It contains no numbered Theorem, Lemma, Proposition, or Corollary asserting membership in \(P\), NP-hardness, parameterized tractability, approximation, or a comparable computational result. Equations (1)–(4) are model definitions, while the Section 5 findings are empirical simulation results. The paper’s claim that it introduces a framework is important, but it is not a named computational result under the programme’s rule.

So there is no paper result I can truthfully claim to cover. Still, the paper supports a plausible *prospective* continuous problem, which I would present as an extension rather than as an anchored mirror.

My lead candidate would be Continuum Counterfactual Dating Evaluation.

Take a finite set \(T\) of complete dating-agent types. A type includes gender and orientation, actual and estimated attractiveness, age, height, BMI, all preference intervals and weights, the truthful profile, the permitted misreporting profiles, the like allowance, the recommendation limit, and the threshold strategy. Since utility depends on the number of previous matches, the current match count is also part of the dynamic state. Thus two agents are the same type only if they are indistinguishable for every future recommendation, liking, matching, and utility calculation.

An instance contains rational type masses \(\mu_t\), with \(\sum_{t\in T}\mu_t=1\), a finite horizon \(R\), per-type like and recommendation quotas, the paper’s preferential matcher, and a finite rational table of compatibility and happiness values obtained from its Equations (3) and (4). The table representation avoids pretending that expressions such as \(x^{0.9}\) have a straightforward exact finite encoding.

For each type \(t\), let \(\mathcal R_t\) be its truthful report together with its permitted deceptive reports. The aggregate misrepresentation decision is

\[ \rho_{t,r}\ge 0 \qquad (r\in\mathcal R_t), \]

where \(\rho_{t,r}\) is the mass of type \(t\) using report \(r\), subject to

\[ \sum_{r\in\mathcal R_t}\rho_{t,r}=\mu_t. \]

The total deceptive mass may be fixed to the paper’s values \(D\in\{0.275,0.813\}\), or supplied as part of the instance.

The matcher then evolves type-level recommendation and matching flows over \(R\) rounds. A flow variable \(z^r_{tu}\) records the mass of evaluator type \(t\) shown candidate type \(u\) in round \(r\); \(y^r_{tu}\) records the resulting mutual-match mass. The flows obey the paper’s recommendation limits, like allowances, no-repeat rule, compatibility-based ranking, and mutual-like condition. The resulting welfare is

\[ W(\rho)=\sum_{r=1}^{R}\sum_{t,u} y^r_{tu} \bigl(U_t(u)+U_u(t)\bigr), \]

with the current-match state included when evaluating \(U\).

The evaluation problem asks for the exact baseline trajectory under truthful reporting, the counterfactual trajectory under \(\rho\), the total welfare difference

\[ \Delta W=W(\rho)-W(\rho_{\mathrm{truth}}), \]

and the welfare change for every type. Its decision form asks whether \(\Delta W\ge \lambda\), or whether a specified type group receives nonnegative change.

This is recognizably connected to the paper: it retains its reported-versus-actual attributes, preferential matchmaking, threshold liking, limited daily likes, repeated rounds, mutual matches, diminishing-return happiness, and truthful/deceitful counterfactual comparison. The mass \(\mu_t\) represents a large cohort of users with identical relevant profiles and strategies. A plausible regime is a national platform with millions of users but a finite catalogue of demographic and preference types, so \(N\) is very large while \(\tau=|T|\ll N\). Clearing denominators in \(\mu\) produces a finite clone population, giving the required high-multiplicity bridge.

For a fixed report plan and an explicit type-level matcher, I would expect the evaluation version to be Class A. Each round is a finite transportation or flow computation over type pairs, with complexity polynomial in \(R\), \(\tau\), and the encoding length. The richer optimization problem—choosing which types misrepresent and which reports they use to maximize welfare—may be difficult because reports change ranking orders and hence the matching graph. Its complexity would be a worthwhile open question, potentially Class C, but it is not a result of this paper.

The regime is plausible but not automatic. The paper itself uses \(8{,}000\) sampled agents and continuously sampled attractiveness, so its experimental populations are not themselves a convincing high-multiplicity limit. The positive case relies on a different scenario: repeated, anonymized user cohorts whose members share all platform-relevant attributes and strategies. The paper’s statement that departing users can be replaced by new agents with similar characteristics helps this interpretation.

The weakest point is that queues and candidate identities matter. Type masses alone do not determine a finite run’s recommendation order or no-repeat history. A legitimate mirror must therefore specify a clone-symmetric mean-field queue or an aggregate recommendation kernel. If that kernel is merely chosen to make the problem tractable, the result becomes a new mean-field model rather than the authors’ dating simulation. Likewise, atomless individual users cannot have meaningful unilateral strategic influence; “strategic misrepresentation” must become a positive-mass type or cohort action.

Thus the best positive conclusion is modest: the paper offers a credible source for a continuous population-level dating-welfare problem, likely tractable in a fixed-kernel finite-type regime. But under the programme’s strict evidence rule, it has no named computational anchor, so it cannot receive a fully supported mirror verdict without relaxing that rule.

The case AGAINST (opponent, writing after the proponent)

The negative case is decisive at the programme’s stated anchoring standard: this paper has no qualifying computational result to continuize. It contains no theorem, lemma, corollary, complexity classification, algorithm, approximation guarantee, or computational decision problem. Equations (1)–(4) define the simulation; the Section 5 claims are empirical outcomes of sampled runs. “We introduce an extensible simulation framework” is a valuable contribution, but it is not a named computational result. Thus the proposed mirror cannot be a mirror of a result in this paper; it is necessarily a new research project.

The proposed Continuum Counterfactual Dating Evaluation also fails as a faithful population continuization. The simulation is not determined by static agent attributes. Each agent has an individual recommendation queue, an individual set of profiles already encountered, an evolving like allowance, a match count \(m_i\), and a counterfactual history. Two agents with identical attractiveness, preferences, reports, and strategy cease to be the same type as soon as they receive different queues or matches. Including the current match count in the type does not solve this: the no-repeat rule requires the identity of every previously encountered candidate, not merely the number of matches.

One can encode those histories as types, but then the type space grows with the population and with the horizon; in a large clone population, initially identical agents rapidly acquire distinct histories. That is precisely the loss of high multiplicity the mirror is supposed to exploit. Alternatively, one can quotient away identities and replace the queue process by an aggregate recommendation kernel. But then the fixed-kernel model is a new mean-field dating model, not a continuization of the paper’s preferential matcher and restricted queues.

There is also a genuine continuum degeneration. The paper gives each user \(O(1)\) recommendations and likes per round. With \(N\) users, each positive-mass type contains \(\Theta(N)\) candidates, while each user sees only \(O(1)\) of them. As \(N\to\infty\), the probability that a user encounters the same individual twice tends to zero, so the paper’s explicit no-repeat mechanism disappears. Preserving it requires scaling recommendation limits with \(N\), or retaining finite candidate pools; either choice changes the regime rather than merely restating it in mass terms.

Finally, \(\rho_{t,r}\) changes the question. In the paper, deceitfulness is assigned exogenously after observing the baseline: the bottom \(X\%\) of realized happiness are made deceptive, and the same sampled population is rerun. If reports are fixed, the proposed problem is an aggregate simulation or flow evaluation, not a computational result from the paper. If reports are optimized, it becomes a new platform-design or strategic-behavior problem absent from the paper. Comparing who benefits additionally requires coupling two individual stochastic trajectories; the marginal distribution \(\mu\) alone cannot identify those paired changes.

A cohort-based mean-field dating model could certainly be worthwhile in its own right. That is the weakness in claiming impossibility “in any scenario.” But it would be an extension inspired by the paper, not a continuous mirror of one of its computational results. Under ChoCo’s rules, every proposed anchor therefore fails.

The adversarial triple: the proponent anchors on up to three named results; the opponent sees that case and must defeat every anchor; the judge decides which case convinced it. These are the pipeline’s own outputs, generated by tools/triple_run.py — no human edited them. The paper’s own text is not reproduced here beyond the quoted statement above.