-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathpaper.html
More file actions
1186 lines (1185 loc) · 52.9 KB
/
Copy pathpaper.html
File metadata and controls
1186 lines (1185 loc) · 52.9 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<meta charset="utf-8" />
<meta name="generator" content="pandoc" />
<meta name="viewport" content="width=device-width, initial-scale=1.0, user-scalable=yes" />
<title>arxiv-paper</title>
<style>
/* Default styles provided by pandoc.
** See https://pandoc.org/MANUAL.html#variables-for-html for config info.
*/
html {
color: #1a1a1a;
background-color: #fdfdfd;
}
body {
margin: 0 auto;
max-width: 36em;
padding-left: 50px;
padding-right: 50px;
padding-top: 50px;
padding-bottom: 50px;
hyphens: auto;
overflow-wrap: break-word;
text-rendering: optimizeLegibility;
font-kerning: normal;
}
@media (max-width: 600px) {
body {
font-size: 0.9em;
padding: 12px;
}
h1 {
font-size: 1.8em;
}
}
@media print {
html {
background-color: white;
}
body {
background-color: transparent;
color: black;
font-size: 12pt;
}
p, h2, h3 {
orphans: 3;
widows: 3;
}
h2, h3, h4 {
page-break-after: avoid;
}
}
p {
margin: 1em 0;
}
a {
color: #1a1a1a;
}
a:visited {
color: #1a1a1a;
}
img {
max-width: 100%;
}
svg {
height: auto;
max-width: 100%;
}
h1, h2, h3, h4, h5, h6 {
margin-top: 1.4em;
}
h5, h6 {
font-size: 1em;
font-style: italic;
}
h6 {
font-weight: normal;
}
ol, ul {
padding-left: 1.7em;
margin-top: 1em;
}
li > ol, li > ul {
margin-top: 0;
}
blockquote {
margin: 1em 0 1em 1.7em;
padding-left: 1em;
border-left: 2px solid #e6e6e6;
color: #606060;
}
code {
font-family: Menlo, Monaco, Consolas, 'Lucida Console', monospace;
font-size: 85%;
margin: 0;
hyphens: manual;
}
pre {
margin: 1em 0;
overflow: auto;
}
pre code {
padding: 0;
overflow: visible;
overflow-wrap: normal;
}
.sourceCode {
background-color: transparent;
overflow: visible;
}
hr {
border: none;
border-top: 1px solid #1a1a1a;
height: 1px;
margin: 1em 0;
}
table {
margin: 1em 0;
border-collapse: collapse;
width: 100%;
overflow-x: auto;
display: block;
font-variant-numeric: lining-nums tabular-nums;
}
table caption {
margin-bottom: 0.75em;
}
tbody {
margin-top: 0.5em;
border-top: 1px solid #1a1a1a;
border-bottom: 1px solid #1a1a1a;
}
th {
border-top: 1px solid #1a1a1a;
padding: 0.25em 0.5em 0.25em 0.5em;
}
td {
padding: 0.125em 0.5em 0.25em 0.5em;
}
header {
margin-bottom: 4em;
text-align: center;
}
#TOC li {
list-style: none;
}
#TOC ul {
padding-left: 1.3em;
}
#TOC > ul {
padding-left: 0;
}
#TOC a:not(:hover) {
text-decoration: none;
}
code{white-space: pre-wrap;}
span.smallcaps{font-variant: small-caps;}
div.columns{display: flex; gap: min(4vw, 1.5em);}
div.column{flex: auto; overflow-x: auto;}
div.hanging-indent{margin-left: 1.5em; text-indent: -1.5em;}
/* The extra [class] is a hack that increases specificity enough to
override a similar rule in reveal.js */
ul.task-list[class]{list-style: none;}
ul.task-list li input[type="checkbox"] {
font-size: inherit;
width: 0.8em;
margin: 0 0.8em 0.2em -1.6em;
vertical-align: middle;
}
.display.math{display: block; text-align: center; margin: 0.5rem auto;}
</style>
</head>
<body>
<h1
id="the-data-trust-index-a-multidimensional-framework-for-evaluating-health-data-integrity-in-ai-systems">The
Data Trust Index: A Multidimensional Framework for Evaluating Health
Data Integrity in AI Systems</h1>
<p><strong>Jason Alan Snyder</strong>¹</p>
<p>¹SuperTruth Inc., United States</p>
<p><strong>Correspondence:</strong> jas@supertruth.ai</p>
<p><strong>Submitted:</strong> April 2026</p>
<hr />
<h2 id="abstract">Abstract</h2>
<p>Health data powering artificial intelligence systems varies widely in
provenance, consent coverage, temporal currency, and clinical
validation, yet most AI pipelines treat data as uniformly trustworthy
once ingested. This paper introduces the Data Trust Index (DTI), a
multidimensional scoring framework that assigns a continuous integrity
score from 0 to 100 to health data records prior to their use in
downstream AI inference. DTI decomposes trust across eight weighted
dimensions: Provenance, Consent, Recency, Quality, Concordance,
Validation, Breadth, and Stability. We describe the theoretical basis
for each dimension, the time-decay functions applied to temporal
signals, and the concordance methodology used to corroborate signals
across independent sources. We present a production implementation case
study with imaware, a direct-to-consumer diagnostics company operating
105,000 diagnostic records, where DTI-based segmentation reduced data
preparation cycles from three weeks to two hours and surfaced a customer
segment accounting for 20% of revenue that was invisible in fragmented
data views. We situate DTI within the regulatory landscape of the 21st
Century Cures Act, the FDA’s AI/ML-Based Software as a Medical Device
(SaMD) guidance, and TEFCA, arguing that a standardized trust metric for
health data is a prerequisite for safely deploying clinical AI at scale.
The DTI framework and its underlying ALDR engine are covered by U.S.
patent application SuperTruth0010CP1.</p>
<p><strong>Keywords:</strong> health data quality, data integrity, AI/ML
in healthcare, data provenance, consent management, clinical decision
support</p>
<p><strong>ACM CCS:</strong> Computing methodologies → Machine learning
→ Machine learning approaches; Applied computing → Life and medical
sciences → Health informatics; Security and privacy → Privacy
protections → Data anonymization and sanitization</p>
<hr />
<h2 id="introduction">1. Introduction</h2>
<p>The dominant assumption in machine learning pipeline design is that
data, once collected and stored, can be treated as fixed input. In
health AI, this assumption is dangerous. A fasting glucose measurement
from a CLIA-certified laboratory with documented chain of custody
carries fundamentally different epistemic weight than a self-reported
blood sugar entry in a consumer wellness app. Yet most clinical AI
systems ingest both without distinction, relying on dataset-level
quality audits rather than record-level trust assessment [1, 2].</p>
<p>The consequences are not hypothetical. The FDA’s 2021 action plan for
AI/ML-Based Software as a Medical Device explicitly identified “good ML
practices” as a gap area, citing the absence of standardized methods for
characterizing training and inference data quality [3]. Research in
clinical NLP has documented that downstream model performance degrades
measurably when training data contains unlabeled missingness patterns,
consent boundary violations, or temporal staleness [4, 5]. In federated
learning environments, where data never leaves its source institution,
the absence of a common quality signal makes model aggregation across
sites an act of statistical faith [6].</p>
<p>The problem is structural. Health data originates across dozens of
modalities (laboratory assays, wearable biometrics, clinical notes,
claims, genomics, SDOH surveys) with incompatible provenance standards,
variable consent frameworks, and modality-specific temporal decay rates.
No existing standard provides a unified, record-level signal that an AI
system can condition on before incorporating a data point into training
or inference.</p>
<p>This paper introduces the Data Trust Index (DTI), a framework that
produces exactly such a signal. DTI scores each health data record on a
0-100 scale by evaluating eight dimensions of data quality and
trustworthiness. The score is composable (dimensions can be reweighted
for specific use cases), interpretable (each dimension contribution is
reported separately), and actionable (score tiers map to recommended use
contexts from exploratory research to regulatory submission).</p>
<p>The remainder of this paper is structured as follows. Section 2
reviews related work in health data quality, data provenance, and trust
metrics. Section 3 defines the eight DTI dimensions and their default
weights. Section 4 describes the scoring methodology, including
time-decay functions and concordance calculations. Section 5 describes
the SuperTruth pipeline implementing DTI at scale. Section 6 presents
the imaware case study. Section 7 discusses regulatory implications.
Section 8 concludes.</p>
<hr />
<h2 id="background-and-related-work">2. Background and Related Work</h2>
<h3 id="health-data-quality-frameworks">2.1 Health Data Quality
Frameworks</h3>
<p>Health data quality has been studied extensively in the context of
electronic health records (EHRs). Weiskopf and Weng [7] proposed a
framework for EHR data quality that identified completeness,
correctness, concordance, plausibility, and currency as key dimensions.
The ONC’s Interoperability Standards Advisory and HL7 FHIR quality
profiles address structural conformance but not epistemic
trustworthiness [8]. The RECORD statement for observational studies and
the EQUATOR network’s reporting guidelines address how data should be
described in publications but do not provide a runtime scoring mechanism
[9].</p>
<p>In the clinical trials domain, ICH E6(R2) Good Clinical Practice
guidelines and FDA 21 CFR Part 11 address data integrity for regulatory
submissions, focusing on audit trails, electronic signatures, and system
validation [10]. These frameworks are rigorous but narrow: they apply to
controlled trial environments, not the heterogeneous data streams
feeding real-world AI systems.</p>
<h3 id="data-provenance">2.2 Data Provenance</h3>
<p>Provenance research in computer science has produced the PROV-DM
standard [11], W3C’s provenance model for linked data, and
domain-specific extensions for scientific workflows [12]. In healthcare,
the HL7 FHIR Provenance resource provides a schema for recording data
origin and transformation history, but adoption is inconsistent and the
resource does not encode a quality signal [13]. Buneman et al.’s
foundational work on “why provenance” [14] distinguishes between why a
data item appears in a result and where it came from, a distinction
directly relevant to health data where source pedigree (CLIA-certified
lab vs. consumer device) matters independently of chain-of-custody
completeness.</p>
<h3 id="consent-and-data-governance">2.3 Consent and Data
Governance</h3>
<p>The HIPAA Privacy Rule [15] and the HITECH Act [16] establish minimum
consent requirements for protected health information, but neither
specifies how consent scope should be encoded in data systems or how
consent completeness should be scored. The 21st Century Cures Act [17]
introduced information blocking provisions and mandated TEFCA as a
national interoperability framework, but TEFCA’s trust framework
addresses organizational participation, not individual record-level
consent quality [18]. GDPR Article 7 [19] requires demonstrable consent
but provides no scoring mechanism. The Common Rule [20] governs research
consent but does not generalize to clinical or commercial contexts.</p>
<h3 id="temporal-data-quality">2.4 Temporal Data Quality</h3>
<p>Temporal aspects of data quality have received attention in database
research [21] and in clinical informatics. Xiao et al. [22] demonstrated
that EHR-based models degrade when trained on features with high
temporal missingness. In wearable data specifically, Bent et al. [23]
showed that heart rate variability features have meaningfully different
stationarity properties over time than discrete biomarker panels,
motivating modality-specific decay modeling.</p>
<h3 id="gap">2.5 Gap</h3>
<p>The frameworks reviewed share a structural limitation: they are
designed for documentation, reporting, or static audit, not for runtime
use. Weiskopf and Weng’s five-dimension framework has been cited over
800 times and represents the dominant approach to EHR data quality
characterization, but it produces no score and provides no mechanism for
conditioning a machine learning pipeline on data quality at inference
time. The RECORD statement and ICH E6 guidelines are publication and
trial-design standards; they cannot be applied to individual records
flowing through a production system. FDA’s SaMD action plan identifies
training data quality as a critical gap but defines no quantitative
standard for addressing it.</p>
<p>The consent gap is particularly acute. HIPAA and GDPR establish legal
requirements for consent but specify neither how to encode consent
completeness in a data record nor how to propagate consent revocation to
downstream consumers. No existing framework provides a runtime,
per-record signal that an AI system can use to confirm that the data it
is incorporating has valid, current, appropriately scoped consent.</p>
<p>The temporal gap is similarly unaddressed. All existing frameworks
treat data quality as a static property measured at a point in time. In
production AI systems that re-use historical data for ongoing inference,
a record’s quality degrades as it ages. No existing framework applies
modality-specific decay functions to produce a recency-adjusted score
that reflects a record’s current trustworthiness rather than its
trustworthiness at the time of collection.</p>
<p>DTI addresses all three gaps simultaneously: it produces a runtime,
per-record score that integrates provenance, consent, temporal currency,
quality, cross-source concordance, clinical validation, dimensional
richness, and signal stability into a single auditable signal usable in
production AI pipelines.</p>
<hr />
<h2 id="the-dti-framework">3. The DTI Framework</h2>
<p>DTI models trustworthiness as a weighted sum of eight dimension
scores, each normalized to 0-100. The default weight vector reflects the
relative evidentiary importance of each dimension for clinical AI
applications, determined through a combination of regulatory guidance
review, clinical expert elicitation, and empirical analysis of feature
sensitivity in downstream model performance.</p>
<h3 id="dimension-definitions">3.1 Dimension Definitions</h3>
<p><strong>Table 1: DTI Dimensions, Weights, and Scoring
Anchors</strong></p>
<table>
<colgroup>
<col style="width: 4%" />
<col style="width: 18%" />
<col style="width: 13%" />
<col style="width: 29%" />
<col style="width: 34%" />
</colgroup>
<thead>
<tr>
<th>#</th>
<th>Dimension</th>
<th>Weight</th>
<th>Low Score (0-40)</th>
<th>High Score (80-100)</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>Provenance</td>
<td>25%</td>
<td>Unknown source, no chain of custody</td>
<td>CLIA/CAP-certified lab, full chain of custody documented</td>
</tr>
<tr>
<td>2</td>
<td>Consent</td>
<td>20%</td>
<td>Implied or absent consent, no revocation mechanism</td>
<td>Explicit, scoped, time-bounded consent with documented revocation
path</td>
</tr>
<tr>
<td>3</td>
<td>Recency</td>
<td>15%</td>
<td>Stale beyond modality-specific threshold</td>
<td>Collected within decay floor for the modality</td>
</tr>
<tr>
<td>4</td>
<td>Quality</td>
<td>10%</td>
<td>High missingness, noise above clinical threshold</td>
<td>Complete fields, noise below floor, no imputation flags</td>
</tr>
<tr>
<td>5</td>
<td>Concordance</td>
<td>10%</td>
<td>Single source, no corroboration</td>
<td>Corroborated across two or more independent sources</td>
</tr>
<tr>
<td>6</td>
<td>Validation</td>
<td>10%</td>
<td>No clinical linkage, no outcome correlation</td>
<td>Peer-reviewed evidence linking biomarker to clinical outcome</td>
</tr>
<tr>
<td>7</td>
<td>Breadth</td>
<td>5%</td>
<td>Single modality, minimal contextual signals</td>
<td>Multi-modal coverage: biomarkers, biometrics, contextual/SDOH</td>
</tr>
<tr>
<td>8</td>
<td>Stability</td>
<td>5%</td>
<td>High intra-individual variance, poor test-retest</td>
<td>Low variance, strong test-retest reliability documented</td>
</tr>
</tbody>
</table>
<p><strong>Total: 100%</strong></p>
<h3 id="dimension-details">3.2 Dimension Details</h3>
<p><strong>Provenance (25%).</strong> Provenance encodes source pedigree
and custody chain. The highest scores require CLIA (Clinical Laboratory
Improvement Amendments) or CAP (College of American Pathologists)
certification for laboratory data, documented instrument calibration,
and an unbroken audit trail from collection to storage.
Consumer-generated data (wearable sensors, self-report) receives
provenance scores conditioned on device FDA clearance status and
documented calibration protocols. Each custody transfer in the chain is
scored for documentation completeness; gaps propagate downward.</p>
<p><strong>Consent (20%).</strong> Consent scoring evaluates four
sub-dimensions: explicitness (implied vs. informed vs. explicit), scope
(how narrowly the consent defines permitted uses), duration (whether the
consent includes an expiration or requires renewal), and revocation
hygiene (whether a documented mechanism exists and has been tested).
Consent tiers map to five levels in SuperTruth’s ConsentOS layer, from
implicit collection through regulatory-grade consent with cryptographic
attestation.</p>
<p><strong>Recency (15%).</strong> Recency applies modality-specific
time-decay functions rather than a uniform staleness threshold. A
complete blood count (CBC) ages more rapidly than a stable genomic
variant; continuous HRV data from a wearable decays differently than an
annual lipid panel. Each modality in the DTI taxonomy has a decay
half-life derived from clinical literature on within-individual
stability. The recency score at time <em>t</em> for a measurement taken
at time <em>t₀</em> is computed as a function of elapsed time and
modality-specific decay parameters, described in Section 4.</p>
<p><strong>Quality (10%).</strong> Quality captures structural
completeness, missingness rate, and signal-to-noise characteristics.
Missingness is scored at the field level; records with imputed values
receive partial credit conditional on the imputation method documented.
Noise floor assessment uses modality-specific clinical reference ranges:
a value three standard deviations outside the reference range without a
documented clinical explanation triggers a quality flag.</p>
<p><strong>Concordance (10%).</strong> Concordance scores the degree to
which a measurement is corroborated by independent sources measuring the
same construct. A self-reported weight that matches a scale measurement
from a clinical visit within 30 days receives high concordance. A
laboratory glucose value that conflicts with a contemporaneous CGM
reading without documented explanation receives low concordance.
Independence of sources is assessed structurally: two measurements from
the same device or the same EHR system are not treated as
independent.</p>
<p><strong>Validation (10%).</strong> Validation scores the clinical
evidence base for the biomarker or measurement type. Biomarkers with
strong peer-reviewed evidence linking them to clinical outcomes (e.g.,
HbA1c and diabetes outcomes, LDL-C and cardiovascular events) receive
high validation scores. Novel or exploratory biomarkers with limited
evidence base receive lower scores. This dimension is a property of the
measurement type, not the individual record, and is updated as the
literature evolves.</p>
<p><strong>Breadth (5%).</strong> Breadth scores dimensional richness:
the number of distinct modalities represented for a given patient or
record set. A record set that includes laboratory panels, biometric
wearable data, clinical notes, and SDOH signals receives higher breadth
scores than a set containing only claims data. Breadth penalizes
over-reliance on any single data modality.</p>
<p><strong>Stability (5%).</strong> Stability scores within-individual
measurement variance and test-retest reliability. High stability scores
require documented low intra-individual coefficient of variation (CV)
for the measurement type and, where available, direct test-retest data.
Stability is particularly important for biomarkers used as monitoring
signals in longitudinal AI applications.</p>
<h3 id="score-tiers">3.3 Score Tiers</h3>
<p>DTI scores map to four operational tiers:</p>
<ul>
<li><strong>Bronze (55-69):</strong> Suitable for exploratory analysis,
hypothesis generation, and non-clinical research.</li>
<li><strong>Silver (70-79):</strong> Suitable for operational analytics,
population health, and commercial AI applications.</li>
<li><strong>Gold (80-89):</strong> Suitable for clinical decision
support and patient-facing applications.</li>
<li><strong>Platinum (90+):</strong> Suitable for regulatory submission,
clinical trial support, and high-stakes AI inference.</li>
</ul>
<h3 id="weight-justification-and-sensitivity">3.4 Weight Justification
and Sensitivity</h3>
<p>The default weight vector was developed through a three-phase
process. First, a regulatory guidance review mapped each dimension to
requirements in HIPAA, 21st Century Cures, FDA SaMD guidance, and ICH
E6(R2), identifying which dimensions are most frequently cited as
failure modes in regulatory enforcement actions and adverse event
reports. Second, a structured elicitation with clinical informaticists
and regulatory specialists rank-ordered dimensions by their estimated
contribution to downstream model failure in clinical AI contexts. Third,
a preliminary sensitivity analysis examined how composite score
orderings responded to single-dimension weight perturbations of ±5
percentage points. Across the dimensions examined, score orderings
showed high stability under moderate weight variation. Prospective,
controlled sensitivity studies across multiple deployment contexts are
included in the future work agenda (Section 7.4).</p>
<p>Provenance receives the highest weight (25%) because source pedigree
errors are the least recoverable class of data quality failure: a record
from an unknown source cannot be remediated by downstream processing.
Consent receives the second-highest weight (20%) because consent
violations carry direct regulatory and legal consequences independent of
data quality. Recency (15%) ranks third because temporal staleness is
the most common cause of silent model degradation in production clinical
AI deployments.</p>
<p>Institutions with domain-specific requirements may rebalance weights
within the constraint that no single dimension exceeds 40% of the
composite, preserving multi-dimensional coverage.</p>
<h3 id="comparison-to-existing-frameworks">3.5 Comparison to Existing
Frameworks</h3>
<p><strong>Table 2: DTI Compared to Existing Health Data Quality
Frameworks</strong></p>
<table style="width:100%;">
<colgroup>
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr>
<th>Framework</th>
<th>Dimensions</th>
<th>Record-Level Score</th>
<th>Runtime Applicable</th>
<th>Regulatory Mapping</th>
<th>Consent Scoring</th>
</tr>
</thead>
<tbody>
<tr>
<td>Weiskopf & Weng (2013) [7]</td>
<td>5 (completeness, correctness, concordance, plausibility,
currency)</td>
<td>No</td>
<td>No</td>
<td>Partial</td>
<td>No</td>
</tr>
<tr>
<td>ONC Interoperability Standards (FHIR/USCDI)</td>
<td>Structural conformance, completeness</td>
<td>No</td>
<td>No</td>
<td>ONC rules, information blocking</td>
<td>No</td>
</tr>
<tr>
<td>TEFCA Trust Framework</td>
<td>Organizational participation</td>
<td>No</td>
<td>No</td>
<td>ONC/TEFCA QHINs</td>
<td>No</td>
</tr>
<tr>
<td>FDA SaMD Action Plan (2021) [3]</td>
<td>Training data documentation</td>
<td>No</td>
<td>No</td>
<td>FDA</td>
<td>No</td>
</tr>
<tr>
<td>ICH E6(R2) GCP [10]</td>
<td>Trial data integrity</td>
<td>No</td>
<td>Trial environments only</td>
<td>FDA/EMA</td>
<td>Yes (narrow)</td>
</tr>
<tr>
<td>RECORD Statement [9]</td>
<td>Reporting completeness</td>
<td>No</td>
<td>No</td>
<td>None</td>
<td>No</td>
</tr>
<tr>
<td><strong>DTI (this work)</strong></td>
<td><strong>8</strong></td>
<td><strong>Yes (0-100)</strong></td>
<td><strong>Yes</strong></td>
<td><strong>FDA, HIPAA, TEFCA, Cures Act</strong></td>
<td><strong>Yes (5 tiers)</strong></td>
</tr>
</tbody>
</table>
<p>The primary differentiator of DTI is its runtime applicability at the
record level. Existing frameworks are designed for documentation,
reporting, or static audit; none produces a score that can be used as an
input feature or filter in a live AI pipeline.</p>
<hr />
<h2 id="scoring-methodology">4. Scoring Methodology</h2>
<h3 id="composite-score">4.1 Composite Score</h3>
<p>The DTI composite score S for record r is computed as:</p>
<p>S(r) = Σᵢ wᵢ · dᵢ(r)</p>
<p>where wᵢ is the weight for dimension i and dᵢ(r) is the normalized
dimension score (0-100) for record r on dimension i. Default weights are
defined in Table 1. For domain-specific deployments (e.g., oncology
trial data where regulatory-grade consent is mandatory), weights can be
rebalanced by the deploying organization within constraints defined in
the SuperTruth configuration layer.</p>
<p><strong>Algorithm 1: DTI Record Scoring</strong></p>
<pre><code>Input: record r, weight vector W, modality m, timestamp t_collect
Output: composite score S(r), per-dimension scores D(r), confidence interval CI(r)
1. P ← score_provenance(r.source, r.custody_chain, r.certification)
2. Co ← score_consent(r.consent_type, r.consent_scope, r.consent_expiry, r.revocation_status)
3. Δt ← current_time − t_collect
4. Re ← recency_decay(Δt, m) // Section 4.2
5. Q ← score_quality(r.fields, r.missingness, r.noise_floor, m)
6. if r.has_independent_source then
7. C ← concordance_score(r, r.independent_source, m) // Section 4.3
8. else
9. C ← 50 // neutral default
10. Va ← lookup_validation_score(m) // literature-derived, not record-specific
11. Br ← score_breadth(r.modalities)
12. St ← score_stability(r.variance_history, m)
13. D(r) ← [P, Co, Re, Q, C, Va, Br, St]
14. S(r) ← Σᵢ W[i] · D(r)[i]
15. CI(r) ← propagate_confidence(D(r), r.documentation_completeness)
16. return S(r), D(r), CI(r)</code></pre>
<p>Lines 6-9 reflect the asymmetric treatment of concordance: the
absence of an independent source is not penalized (defaulting to neutral
50) rather than scored negatively, because many legitimate data sources
have no contemporaneous independent measurement available.</p>
<h3 id="recency-decay-functions">4.2 Recency Decay Functions</h3>
<p>For each modality m with decay half-life τₘ (in days), the recency
score at elapsed time Δt is:</p>
<p>R(Δt, m) = 100 · exp(−λₘ · Δt)</p>
<p>where λₘ = ln(2) / τₘ. Scores are floored at 0. The decay function is
applied from the timestamp of collection, not receipt, where collection
timestamp is documented.</p>
<p><strong>Table 2: Modality-Specific Recency Decay
Half-Lives</strong></p>
<table>
<colgroup>
<col style="width: 33%" />
<col style="width: 33%" />
<col style="width: 33%" />
</colgroup>
<thead>
<tr>
<th>Modality</th>
<th>Half-Life τ (days)</th>
<th>Clinical Basis</th>
</tr>
</thead>
<tbody>
<tr>
<td>Complete blood count (CBC)</td>
<td>90</td>
<td>Erythrocyte and leukocyte turnover rates</td>
</tr>
<tr>
<td>Lipid panel</td>
<td>180</td>
<td>Cardiovascular risk signal stability</td>
</tr>
<tr>
<td>HbA1c</td>
<td>90</td>
<td>Reflects 3-month erythrocyte lifecycle</td>
</tr>
<tr>
<td>Continuous HRV aggregates</td>
<td>14</td>
<td>High intra-individual variability; short stationarity window</td>
</tr>
<tr>
<td>Fasting glucose (discrete)</td>
<td>30</td>
<td>Metabolic state changes on weekly-monthly timescale</td>
</tr>
<tr>
<td>SDOH survey data</td>
<td>365</td>
<td>Social determinants change slowly relative to biomarkers</td>
</tr>
<tr>
<td>Clinical notes (structured fields)</td>
<td>180</td>
<td>Diagnosis and medication states are semi-stable</td>
</tr>
<tr>
<td>Genomic variants</td>
<td>∞</td>
<td>Constitutional variants do not change; scored on collection quality
only</td>
</tr>
<tr>
<td>Consumer wearable daily aggregates</td>
<td>30</td>
<td>Device accuracy degrades; behavioral patterns shift</td>
</tr>
</tbody>
</table>
<p>Half-life values are derived from clinical literature on
within-individual measurement stability and represent the elapsed time
at which the recency score falls to 50 (i.e., the measurement retains
half its original recency credit). Organizations may calibrate
half-lives to their specific clinical context.</p>
<h3 id="concordance-calculation">4.3 Concordance Calculation</h3>
<p>For a target measurement m from source s₁, concordance with
independent measurement m’ from source s₂ at time proximity δt is:</p>
<p>C(m, m’) = max(0, 100 · (1 − |m − m’| / tolerance(type)))</p>
<p>where tolerance(type) is the clinically acceptable measurement
variation for the measurement type. If no contemporaneous independent
measurement exists, concordance defaults to 50 (neutral, not penalized).
Concordance scores above 85 for a measurement pair elevate the composite
score; scores below 30 trigger a concordance flag that is returned to
the consuming application.</p>
<p><strong>Example.</strong> A fasting glucose measurement of 110 mg/dL
from an at-home collection kit is compared against a clinical laboratory
value of 113 mg/dL drawn four days later. The clinically accepted
measurement tolerance for fasting glucose from point-of-care devices is
±15 mg/dL per ISO 15197:2013 [24] and FDA 21 CFR 862.1345. The
concordance score is:</p>
<p>C = max(0, 100 · (1 − |110 − 113| / 15)) = max(0, 100 · (1 − 0.20)) =
80</p>
<p>This score (80) indicates good concordance. The two measurements
agree within clinical tolerance despite originating from independent
collection events and different laboratory systems. If the at-home value
had been 130 mg/dL (a 20 mg/dL discrepancy exceeding the 15 mg/dL
tolerance), the concordance score would be max(0, 100 · (1 − 1.33)) = 0,
triggering a concordance flag for clinical review.</p>
<h3 id="confidence-intervals-and-propagation">4.4 Confidence Intervals
and Propagation</h3>
<p>Each dimension score includes a confidence interval derived from data
completeness. When provenance documentation is partial, the provenance
score carries a wider interval. The composite score’s confidence
interval is propagated from dimension-level intervals using standard
error propagation. Consuming applications can filter records by
composite score confidence in addition to point estimate.</p>
<hr />
<h2 id="implementation-the-supertruth-pipeline">5. Implementation: The
SuperTruth Pipeline</h2>
<p>SuperTruth implements DTI through a four-stage pipeline illustrated
in Figure 1.</p>
<pre><code> Raw Data Sources Consuming Applications
┌─────────────┐ ┌──────────────────────┐
│ Clinical EHR│──┐ │ AI Models / LLMs │
│ Laboratory │──┤ │ Intelligence Portal │
│ Wearable │──┼──► IntegrityNet ──► DTI Engine ──► ConsentOS ──► Data Reservoir ──►│ Federated Networks │
│ Claims/Payer│──┤ (Ingest, │ (Score │ (Consent │ (Scored, │ MyBio.Health │
│ Patient SDOH│──┘ Zero-Copy) │ 0-100) │ Governance) │ Query-Ready) │ Research Portals │
└─────────────┘ └──────────────────────┘</code></pre>
<p><em>Figure 1: The SuperTruth four-stage pipeline. Data from
heterogeneous sources enters IntegrityNet via zero-copy connectors. Each
record is scored by the DTI Engine across eight dimensions. ConsentOS
enforces consent boundaries. The Data Reservoir exposes scored,
consent-governed data to consuming applications. DTI score travels
permanently with every record through all stages.</em></p>
<p><strong>Table 3: Pipeline Stage Summary</strong></p>
<table>
<colgroup>
<col style="width: 20%" />
<col style="width: 20%" />
<col style="width: 20%" />
<col style="width: 20%" />
<col style="width: 20%" />
</colgroup>
<thead>
<tr>
<th>Stage</th>
<th>Input</th>
<th>Operation</th>
<th>Output</th>
<th>Patent Coverage</th>
</tr>
</thead>
<tbody>
<tr>
<td>IntegrityNet</td>
<td>Raw records (any source, any format)</td>
<td>Schema normalization, provenance capture, audit trail creation</td>
<td>Structured records with provenance metadata</td>
<td>SuperTruth0010CP1</td>
</tr>
<tr>
<td>DTI Engine</td>
<td>Provenance-annotated records</td>
<td>Eight-dimension scoring via ALDR (Algorithm 1)</td>
<td>Records with DTI score + per-dimension breakdown + confidence
interval</td>
<td>SuperTruth0010CP1</td>
</tr>
<tr>
<td>ConsentOS</td>
<td>Scored records</td>
<td>Consent state enforcement, cryptographic attestation, revocation
propagation</td>
<td>Consent-governed records with use-authorization flags</td>
<td>SuperTruth0010CP1</td>
</tr>
<tr>
<td>Data Reservoir</td>
<td>Scored, consent-governed records</td>
<td>Score-indexed storage, query interface, tier-based access
control</td>
<td>Queryable data layer with DTI as first-class filter</td>
<td>SuperTruth0010CP1</td>
</tr>
</tbody>
</table>
<p>SuperTruth implements DTI through a four-stage pipeline, described
below:</p>
<p><strong>Stage 1: IntegrityNet (Ingestion).</strong> Data is ingested
via zero-copy connectors to source systems: FHIR R4 APIs, HL7 v2 feeds,
SFTP drops from laboratory systems, and direct API integrations with
wearable platforms. Zero-copy means source data is never moved to a
SuperTruth data store; instead, a metadata record describing the data
and its provenance is created. This architecture eliminates a class of
consent and security risk associated with data replication.</p>
<p><strong>Stage 2: DTI Engine (Scoring).</strong> The DTI Engine
evaluates each ingested record against the eight dimensions. The engine
is implemented using the ALDR (Adaptive Learning Data Resilience)
architecture, which comprises three components: MEA (Matrix Exponential
Attention) for cross-source signal alignment, SADC (Self-Adjusting Decay
Control) for modality-specific temporal scoring, and GLMN
(Graph-Laplacian Memory Networks) for longitudinal stability assessment.
ALDR is the subject of U.S. patent application SuperTruth0010CP1.</p>
<p><strong>Stage 3: ConsentOS (Consent Governance).</strong> ConsentOS
enforces consent boundaries on scored data through five additive patient
consent actions: Grant, Modify, Pause, Resume, and Revoke. Each action
is recorded with a cryptographic attestation. Revocation propagates to
all downstream systems in real time, ensuring that a patient’s
withdrawal of consent is honored across the entire pipeline without
manual intervention. ConsentOS maintains an immutable audit log of every
consent state transition.</p>
<p><strong>Stage 4: Data Reservoir (Query-Ready Output).</strong>
Scored, consent-governed data is exposed to consuming applications
through a query interface that includes DTI score as a first-class
filter. Applications can query for all records above a DTI threshold,
within a score tier, or filtered by individual dimension scores. The
reservoir is read-only from the perspective of data consumers; writes
flow only through IntegrityNet.</p>
<hr />
<h2 id="retrospective-deployment-study-imaware">6. Retrospective
Deployment Study: imaware</h2>
<p>This section presents results from a retrospective before/after
deployment study rather than a controlled experiment. The findings are
reported as quantitative operational outcomes of a production DTI
implementation. Controlled replication across multiple institutions is
required before generalizing these findings.</p>
<p>imaware is a direct-to-consumer laboratory diagnostics company
offering at-home collection for panels including cardiometabolic,
thyroid, hormonal, and inflammatory markers. Prior to DTI
implementation, imaware managed approximately 105,000 diagnostic records
across fragmented data systems with no unified quality signal.</p>
<h3 id="study-design">6.0 Study Design</h3>
<p><strong>Setting:</strong> Single-organization retrospective
before/after comparison. <strong>Data:</strong> 105,000 diagnostic
records across imaware’s laboratory information system.
<strong>Intervention:</strong> Full DTI implementation via SuperTruth
IntegrityNet integration. <strong>Comparison period:</strong> 6-month
pre-implementation baseline vs. 6-month post-implementation observation.
<strong>Primary outcomes:</strong> (1) data preparation cycle time; (2)
analyst hours allocated to data quality tasks; (3) identifiable customer
segment structure. <strong>Limitations:</strong> No randomization, no
control arm, single organization. Confounders including organizational
learning effects and concurrent infrastructure improvements cannot be
excluded.</p>
<h3 id="implementation">6.1 Implementation</h3>
<p>SuperTruth integrated with imaware’s laboratory information system
(LIS) via structured API integration with the CLIA-certified processing
laboratory. DTI scoring was applied to all records at ingestion. The
integration required no data movement: IntegrityNet’s zero-copy
connectors maintained data in imaware’s existing systems while appending
DTI metadata.</p>
<h3 id="results">6.2 Results</h3>
<p><strong>Outcome 1: Data preparation cycle time.</strong></p>
<table>
<colgroup>
<col style="width: 25%" />
<col style="width: 25%" />
<col style="width: 25%" />
<col style="width: 25%" />
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Pre-DTI</th>
<th>Post-DTI</th>
<th>Change</th>
</tr>
</thead>
<tbody>
<tr>
<td>Data preparation cycle time</td>
<td>~3 weeks</td>
<td>~2 hours</td>
<td>−95%</td>
</tr>
<tr>
<td>Analyst hours/month on data quality</td>
<td>~200 hours</td>
<td>~0 hours</td>
<td>−200 hrs/month</td>
</tr>
</tbody>
</table>
<p>Prior to DTI, preparing a data set for analysis required
approximately three weeks of manual review, deduplication, and quality
assessment. Post-DTI, the same preparation cycle completed in
approximately two hours using DTI score thresholds to filter records
programmatically, a 95% reduction. Approximately 200 analyst hours per
month previously allocated to data hygiene tasks were freed for
higher-value work.</p>
<p><strong>Outcome 2: Latent segment identification.</strong></p>
<p>DTI-based segmentation of the 105,000-record dataset identified a
customer segment that had been indistinguishable in the pre-DTI
fragmented data view. This segment accounted for 20% of imaware’s
revenue. The segment was characterized by consistent concordance between
at-home and clinical laboratory results, high Recency scores (frequent
re-testing behavior), and multi-panel Breadth. These properties were
measurable only after DTI scoring unified and ranked the underlying
records into a queryable structure.</p>
<p><strong>Qualitative assessment.</strong> Brodie Flanders, CEO of
imaware, stated: “The lab industry has never had a trust standard. DTI
created one.”</p>
<h3 id="prospective-validation-protocol">6.3 Prospective Validation
Protocol</h3>
<p>The following validation study design is proposed to assess DTI’s
impact on downstream model performance across independent institutions.
This study has not yet been conducted; the protocol is published here to
establish a pre-registered methodological framework.</p>
<p><strong>Study objective:</strong> To determine whether training
clinical AI models on DTI-filtered data (records above a specified score
threshold) produces measurably better performance than training on
unfiltered data from the same source population.</p>
<p><strong>Data sources:</strong> Three independent health systems
contributing de-identified records under IRB approval, targeting a
minimum of 50,000 records per site across at least two modalities
(laboratory + wearable or laboratory + clinical notes).</p>
<p><strong>Experimental conditions:</strong> For each site, three model
training conditions: (1) unfiltered baseline (all records), (2) DTI
threshold ≥70 (Silver-tier filter), (3) DTI threshold ≥80 (Gold-tier
filter).</p>
<p><strong>Outcome metrics:</strong> Model calibration (Brier score),
discrimination (AUROC), and positive predictive value on two clinically
validated prediction tasks: 30-day readmission and HbA1c trajectory.</p>
<p><strong>Hypothesis:</strong> Models trained on DTI-filtered data will
show statistically significant improvement in Brier score relative to
unfiltered baseline, with the largest improvement in the Gold-tier
condition. We further hypothesize that improvement will be greater for
sites with higher pre-DTI data fragmentation.</p>
<p><strong>Planned analysis:</strong> Mixed-effects model with site as a
random effect. Significance threshold p < 0.05 with Bonferroni
correction for multiple comparisons. Pre-registration on OSF prior to
data collection.</p>
<h3 id="limitations">6.4 Limitations</h3>
<p>The imaware implementation covers laboratory diagnostics data only;
breadth scores for this data set are accordingly lower than for
multi-modal datasets. The concordance dimension benefited from imaware’s
dual-source architecture (at-home collection followed by clinical
confirmation), which is not universal across laboratory operators.</p>
<p>The reported time reduction from three weeks to two hours reflects a
before/after comparison within a single organization, not a controlled
experiment with a held-out condition. Confounding factors — including
organizational learning effects and concurrent data infrastructure
improvements — cannot be fully excluded. The figure should be treated as
an operational benchmark, not a causal estimate.</p>
<p>The revenue intelligence finding (20% of revenue from a previously
invisible segment) was surfaced post-hoc through DTI-enabled
segmentation. Prospective replication studies are required to establish
whether similar latent segments exist across other diagnostics operators
and whether DTI-based segmentation reliably surfaces them.</p>
<hr />
<h2 id="discussion">7. Discussion</h2>
<h3 id="regulatory-context">7.1 Regulatory Context</h3>
<p>The 21st Century Cures Act [17] mandated open APIs for health data
access and defined information blocking prohibitions, accelerating the
volume of health data available to AI systems without establishing
quality standards for that data. TEFCA [18], implemented by the Office
of the National Coordinator for Health IT, defines a trust framework for