dorsal/arxiv
View SchemaSemiparametric Efficient Data Integration Using the Dual-Frame Sampling Framework
| Authors | Kosuke Morikawa, Jae Kwang Kim |
|---|---|
| Categories | |
| ArXiv ID | 2601.08707vv1 |
| URL | https://arxiv.org/abs/2601.08707 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the non-probability inclusion probability parametrically and attains the semiparametric efficiency bound. We introduce an identifiability condition based on strong monotonicity that identifies sampling-model parameters without instrumental variables, even under informative (non-ignorable) selection, using auxiliary information from the probability sample; it remains valid without record linkage between samples. The second estimator, motivated by a two-stage sampling approximation, avoids explicit modeling of the non-probability mechanism; though not fully efficient, it is efficient within a restricted augmentation class and is robust to misspecification. Simulations and an application to the Culture and Community in a Time of Crisis public simulation dataset show efficiency gains under correct specification and stable performance under misspecification and weak identification. Methods are implemented in the R package \texttt{dfSEDI}.
{
"annotation_id": "e79bc4c4-5a7f-4fac-b7bb-4d260ec9f1f7",
"date_created": "2026-02-17T05:53:16.291000Z",
"date_modified": "2026-02-17T05:53:16.291000Z",
"file_hash": "bb98dc5cd05a476e4f2ba83a71a0752d8ce7c92c24a9b8d3191ce57731796624",
"private": false,
"record": {
"abstract": "Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame data integration and propose two complementary estimators. The first models the non-probability inclusion probability parametrically and attains the semiparametric efficiency bound. We introduce an identifiability condition based on strong monotonicity that identifies sampling-model parameters without instrumental variables, even under informative (non-ignorable) selection, using auxiliary information from the probability sample; it remains valid without record linkage between samples. The second estimator, motivated by a two-stage sampling approximation, avoids explicit modeling of the non-probability mechanism; though not fully efficient, it is efficient within a restricted augmentation class and is robust to misspecification. Simulations and an application to the Culture and Community in a Time of Crisis public simulation dataset show efficiency gains under correct specification and stable performance under misspecification and weak identification. Methods are implemented in the R package \\texttt{dfSEDI}.",
"arxiv_id": "2601.08707",
"authors": [
"Kosuke Morikawa",
"Jae Kwang Kim"
],
"categories": [
"stat.ME",
"math.ST",
"stat.TH"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Semiparametric Efficient Data Integration Using the Dual-Frame Sampling Framework",
"url": "https://arxiv.org/abs/2601.08707",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e2267324-5615-4cea-8eb3-c12a1e781b50",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}