Loan Origination Data Quality Decides Coverage, Not Accuracy
Published lift research favours the model, not the data. Where loan origination data quality decides more, and the three numbers that would settle it.
Alfred BEditorial Reviews
Two research teams have run this comparison directly, and both times the model won. The technique added more predictive accuracy than the extra data did, by a wide margin. Any claim that the pipeline is the better investment starts there.
Loan origination data quality does not beat model choice on predictive lift. Published head-to-head comparisons find the opposite. What the pipeline decides is which applicants reach a model at all, how old their information is when it gets there, and whether a decision can be made without asking again.
Does better data actually beat a better model?
In published comparisons, the modelling technique has moved predictive accuracy more than the added data. FinRegLab's July 2025 study found machine learning raised ROC-AUC by about 2 percent over logistic regression, while adding cash-flow data to credit bureau data added roughly 0.3 to 0.7 percent.
FinRegLab is a non-profit research organisation, and its numbers are not an outlier. Leonardo Gambacorta, Yiping Huang, Han Qiu and Jingyi Wang, in Bank for International Settlements Working Paper 834 of December 2019, decomposed a Chinese fintech lender's credit score and found non-traditional data added 2.2 percent of AUROC against 5.3 from machine learning.
The two samples are nothing alike. FinRegLab's is rich and it says so: it "includes relatively few consumers whose credit bureau data is so limited that they may have a difficult time being scored." The BIS baseline is the opposite, an AUROC of 0.594 on loans to Chinese online merchants, and machine learning still contributed more there.
Thin files are still where added data earns most, on narrower evidence than usually implied. Berg, Burg, Gombović and Puri, in NBER Working Paper 24551 of April 2018, found that among the 6 percent of their German e-commerce sample with no bureau score, digital-footprint variables alone reached an AUC of 72.2 percent, above the 69.6 they reached on scorable customers.
That is a finding about who can be scored, not how finely. It does not rescue the pipeline argument on accuracy, since both head-to-heads put the technique ahead and one ran on a coin-toss baseline. Where cash-flow and alternative data win is on people a bureau file cannot rank.
What does loan origination data quality actually govern?
Loan origination data quality governs coverage more than ranking. The Consumer Financial Protection Bureau reported in May 2015 that, as of December 2010, 26 million American adults had no credit record and another 19 million held records that could not be scored, 9.6 million of those because the file had gone stale.
Input accuracy is a separate problem. The Federal Trade Commission, reporting on 11 February 2013 on 2,968 credit reports, found one in five consumers had an error on at least one, and about one in twenty saw a score change above 25 points once corrected. That swing does not degrade a model. It ranks a real applicant against wrong facts, which no performance statistic surfaces.
Published effect sizes, sorted by what they move. The accuracy levers and the coverage levers use different units.
| Lever | What was measured | Size of effect | Source and date |
|---|---|---|---|
| Machine learning instead of logistic regression, bureau data | ROC-AUC | about +2% | FinRegLab, 1 July 2025 |
| Machine learning instead of a scorecard, fintech data | AUROC | +5.3% | BIS Working Paper 834, December 2019 |
| Adding cash-flow data to credit bureau data | ROC-AUC | +0.3% to +0.7% | FinRegLab, 1 July 2025 |
| Adding non-traditional data to traditional data | AUROC | +2.2% | BIS Working Paper 834, December 2019 |
| Digital footprint instead of bureau score, scorable German e-commerce customers | AUC | 69.6% vs 68.3% | Berg, Burg, Gombović and Puri, NBER WP 24551, April 2018 |
| Digital footprint added to bureau score | AUC | 68.3% to 73.6% | Berg, Burg, Gombović and Puri, NBER WP 24551, April 2018 |
| Digital footprint alone, customers with no bureau score | AUC | 72.2% | Berg, Burg, Gombović and Puri, NBER WP 24551, April 2018 |
| Adults with no credit record, or an unscorable one | share of US adults, December 2010 | 11% plus 8.3% | CFPB Data Point: Credit Invisibles, May 2015 |
| Credit report errors material enough to move a score | share of consumers | about 1 in 20 above 25 points | Federal Trade Commission, 11 February 2013 |
Units differ down the size-of-effect column. The BIS figures are relative percentage change in AUROC, its arithmetic given as (0.607 - 0.5939) / 0.5939 and (0.6391 - 0.607) / 0.607; the FinRegLab figures are changes in the metric itself. In AUROC points the BIS result is plus 1.3 for data against plus 3.2 for model.
Why is cleaning the data you already hold an unreliable fix?
Cleaning existing data is a poor way to buy accuracy. The CleanML benchmark, published at the IEEE International Conference on Data Engineering in 2021, tested five error types across 14 real datasets and seven model families, and found that cleaning outliers had no significant effect in 61 percent of experiments.
The rest is equally deflating, on uneven denominators: 560 outlier experiments, 294 missing-value, 112 duplicate. Cleaning duplicates was insignificant in 67 percent of those and negative in 22. Imputation helped in 49 percent and hurt in 24. The distinction that matters is between scrubbing data after it arrives and changing where it comes from, and only the second is a pipeline decision.
Why do lenders put the money in models?
Model work gets funded because it can be attributed. A model change has an owner, a version, a control group and a number that moves inside a quarter. Pipeline work crosses sales, operations, credit and engineering, and its benefit arrives as files that never became problems.
This is not a lapse of judgement, and the research says so. Nithya Sambasivan and colleagues at Google Research, in a CHI 2021 study of 53 AI practitioners, found 92 percent had hit a data cascade, "opaque in diagnosis and manifestation," and 32.1 percent traced one to conflicting reward systems. The authors put it plainly: "Care of, and improvements to data are not easily 'tracked' or rewarded, as opposed to models."
Detection is genuinely hard, not merely neglected. D. Sculley and colleagues, in the NIPS 2015 paper on hidden technical debt in machine learning systems, wrote that "only a small fraction of real-world ML systems is composed of the ML code," and that data dependencies "may be more difficult to detect" than code dependencies. Establishing how often anomalies actually fire took Eric Breck and colleagues, in the SysML 2019 data validation paper, more than 700 production pipelines at Google.
Supervisors report the same shape. The Bank of England and Financial Conduct Authority survey of 118 firms, published 21 November 2024, found four of the top five perceived current AI risks were data-related, data quality second, and the Office of the Superintendent of Financial Institutions and Financial Consumer Agency of Canada report of 24 September 2024 puts data governance top of the Canadian list.
A two percent ROC-AUC gain on a population you already reach is worth less to most lenders than a several-point gain in the share of applicants who reach a decision at all, and almost nobody measures the second one, so the comparison never happens and the model wins on walkover.
How would a lender sequence the fix?
The sequence that makes pipeline work attributable starts with measurement, not rebuild. Three numbers put the pipeline on the model's footing: the share of applications that reach a decision, the age of each decisive input at decision time, and the share of inputs that came from the issuing institution.
- Count applications that never reach a decision, by cause: abandoned, incomplete, unverifiable, waiting on a third party. That is the denominator the lift research is silent about.
- Time-stamp inputs at decision time, not collection time. A pay stub dated eleven weeks ago and a connection pulled this morning wear the same label in most systems.
- Record provenance as a field: applicant-supplied, employer-supplied, institution-supplied. A model cannot recover it from values alone.
- Report all three beside model performance, so both compete for one budget.
Adjacent evidence points the same way. Federal Reserve Bank of New York Staff Report 836 of February 2018 found technology-based mortgage lenders about 20 percent faster with defaults about 25 percent lower on FHA loans, pointing at process rather than screening. We went through it in the speed and defaults evidence.
None of this argues against a better model. The two investments are measured in different units, and only one has a unit.
What we couldn't verify
Three things here have no traceable source. The cost-of-bad-data figures that circulate in this argument, versions of "$12.9 million a year for the average organisation" and "$3.1 trillion a year to the US economy," appear constantly with no published methodology behind either, and neither is used above.
Any study isolating input timeliness from input source in credit model performance. Nothing found holds the source constant and varies only age.
Any Canadian figure for applications that never reach a decision on missing or unverifiable inputs. The lift research above is American, German and Chinese.
Common questions
Does data quality matter more than model choice in credit underwriting?
Not on predictive accuracy. FinRegLab's July 2025 study found machine learning added about 2 percent ROC-AUC while cash-flow data added 0.3 to 0.7 percent. BIS Working Paper 834 of December 2019 found 5.3 percent from machine learning against 2.2 from non-traditional data.
Where does loan origination data quality make the larger difference?
In coverage and timeliness. The Consumer Financial Protection Bureau reported in May 2015 that, as of December 2010, 26 million US adults had no credit record and 19 million more were unscorable, 9.6 million because the record had gone stale. Those applicants reach no model at all.
Does cleaning data improve model performance?
Unreliably. The CleanML benchmark, published at the IEEE International Conference on Data Engineering in 2021 across 14 datasets and seven model families, found outlier cleaning insignificant in 61 percent of 560 experiments, duplicate cleaning insignificant in 67 percent of 112, and imputation negative in 24 percent of 294.
Why do teams invest in models rather than pipelines?
Because model work is attributable and pipeline work is not. Sambasivan and colleagues at Google Research, in a CHI 2021 study of 53 AI practitioners, found 92 percent had experienced data cascades, and wrote that improvements to data "are not easily 'tracked' or rewarded, as opposed to models."
How much of a machine learning system is actually the model?
A small share. D. Sculley and colleagues, in the NIPS 2015 paper on hidden technical debt in machine learning systems, wrote that "only a small fraction of real-world ML systems is composed of the ML code," and described the surrounding infrastructure as vast and complex.
Carousel works on the input side of this. See how verification fits your flow


