The comparison that flips: a trap in measuring what drives adoption¶
A plain-language companion to the draft working paper "Why Driver Weights Differ: A Cost-Based Microfoundation for Segmented Technology Acceptance" (v0.2, July 2026). The paper carries the model, the proofs, and the verification; this text carries the ideas.
Read the full draft working paper (PDF)
The short version. Suppose you survey organizations about adopting open source and want to compare groups: does ease of adoption matter more to unskilled adopters than to skilled ones? The obvious method (estimate the effect in each group, compare the sizes) can return exactly the wrong answer: it can conclude that effort matters least precisely where it binds most. The paper shows why with a small economic model, quantifies the trap, and states the fix. On random test configurations, the naive comparison misorders the groups 63 percent of the time under one standard statistical setup and 46 percent under another, while comparing ratios of effects within each group gets the ordering right every time. This is a methods paper: no survey data anywhere, and every claim is verified to machine precision.
Where the trap comes from¶
Model adoption the way economists model any yes-or-no decision: an organization adopts when the benefits beat the costs, plus some noise. A group's capability is its unit cost of turning adoption effort into working systems: skilled groups convert cheaply, unskilled ones expensively. In this world, every "driver weight" a survey estimates is a product of two things: how much the driver truly matters to that group (the structural part) and how close the group sits to the tipping point of its adoption decision (the location part). A group near fifty-fifty responds visibly to every nudge; a group that almost never adopts, or almost always does, barely registers any.
The location part is what poisons naive comparisons. Take an unskilled group far from adopting, at a five percent baseline. Effort is its biggest barrier, four times more binding than for the skilled group in the paper's worked example. Yet its measured effort effect comes out smaller (0.095 against 0.139), because the group sits in the flat tail of its adoption curve where nothing moves the needle much. Compare the raw effects and you conclude effort matters least for the least capable, which is backwards.
The fix¶
Divide before you compare. Within each group, take the ratio of two effects, say effort against performance: the location part sits in both and cancels, leaving pure structure. The paper proves that ordering groups by these ratios recovers the true capability ordering under any differences in baseline adoption, and the numerical verification bears it out: ratio comparisons reproduce the true ordering on every sampled configuration, and the exact condition for when raw-level comparisons flip is characterized and then confirmed, with no exceptions, on 200,000 random configurations per statistical setup. Comparing groups at matched baseline adoption rates works too, and eliminates the reversals entirely.
There is a second version of the same trap, aimed at standard survey practice. Multi-group studies often compare standardized coefficients across groups. Standardization bakes each group's variances into the coefficient, so two groups with identical structure but different spread come out looking structurally different. A skilled group in which adoption effort is uniformly low has little effort variance, its standardized effort path shrinks, and the reading becomes "effort does not matter for the capable" when what happened is that the capable do not vary in it. In the paper's checks this reverses a true ordering 17 percent of the time, while unstandardized comparisons on comparable measurement scales never do.
Why this matters beyond one survey¶
The paper is the theory companion of a survey design that asks whether open-source adoption drivers differ across capability segments, and it delivers that project two things: a derivation of its central hypothesis from a cost primitive instead of an assertion, and a set of binding rules for the test (state conclusions on ratios or at matched baselines, report every group's baseline next to any comparison, never compare standardized coefficients across groups). The survey's registered analysis plan adopts the rules wholesale.
The warning travels further than one project. Comparing driver weights across groups is routine in technology-acceptance research, and the components of the problem are known in econometrics: the paper cites the canonical source and claims the packaging, a pre-registrable discipline for acceptance studies, as its contribution. A documented sweep of possible antecedents found partial precedents for the ingredients, and the paper's next revision repositions its claims against that record.
The small print¶
The percentages above (63, 46, 17) describe how often the traps spring across a sampled space of random test configurations; they measure how generic the problem is within that space, and they are no estimate of how often real published studies err. The traps themselves are not a discovery of this paper: sociologists and econometricians have warned for twenty-five years that groups are hard to compare in models of this kind, and the ratio fix rests on a textbook identification fact. What the paper adds is the derivation of the segment differences from a cost primitive, the exact condition under which the ordering flips, the measured frequency, and the import of the discipline into acceptance research, which had not picked it up. All results are theorems about the model, each with a written proof, a numeric cross-check against brute-force computation, and a Monte-Carlo robustness run. No empirical claim about actual adopters is made anywhere in the paper.