Alternative Data That Actually Works (and the Kind That Doesn't)
Nine categories of alternative data graded on one evidence standard. What bank, education and utility data actually prove, and where the research stops.
Alfred BEditorial Reviews
Nine categories get sold under one label, and the evidence is uneven in ways the marketing does not track. Seven have been measured against observed repayment. Two have not, including one everybody treats as settled. Only one has been measured twice by a party with nothing to sell.
Alternative data underwriting works when a signal has measured lift on a lending population, holds up out of time, can be explained to the applicant, and measures the applicant rather than the group. No category here clears all four on published evidence. Bank transaction data comes closest, at under a third of an ROC-AUC point over a bureau file.
What separates a real signal from a sold one?
Four tests. Measured lift on a population that actually borrowed and either repaid or did not. Stability when the model is trained on one period and scored on a later one. An explanation an underwriter can give out loud. And evidence that the variable measures the applicant rather than the neighbourhood, cohort or device fleet the applicant belongs to.
The same four apply to every row below, including the rows this piece would rather grade kindly. Non-isolation counts against a category whether or not the category is fashionable.
Stability has the least evidence behind it, and the one published result is not encouraging. Björkegren and Grissen, in World Bank Policy Research Working Paper 9074 of December 2019, scored mobile phone metadata against repayment for 7,068 borrowers, reaching an AUC of 0.760 on held-out folds inside the sample period and 0.595 when trained early and tested late, against a bureau baseline of 0.565.
The proxy test is where the list gets awkward. Fuster, Goldsmith-Pinkham, Ramadorai and Walther, in the Journal of Finance in 2022, found Black and Hispanic borrowers less likely to gain from machine learning on US mortgage data, driven by statistical flexibility rather than proxy variables. Blattner and Nelson, in a Stanford paper of May 2021, traced the extra noise in scores for under-served groups to thin data. More information about the person, then, rather than about people like the person.
Which alternative data underwriting signals have real published lift?
Bank transaction data, education and employment attributes, and utility or telecom payment history. Each has been measured against observed repayment on a lending population by researchers who were not selling the data. None of the three is cleanly isolated from other inputs, and one rests on borrower behaviour recorded sixteen years ago.
Bank transaction data has the most consistent evidence and the smallest measured gain. FinRegLab, a non-profit research organisation, reported in July 2019 that standing alone, cash-flow metrics generally performed as well as traditional credit scores, among populations and products like the ones studied, where traditional credit history is not available or reliable. The six participants were non-bank lenders to exactly that population.
FinRegLab's July 2025 study is the concession this category lives with. On an out-of-time sample of 424,546 consumers, adding cash-flow data to bureau data moved ROC-AUC from 0.8653 to 0.8682 under logistic regression and from 0.8831 to 0.8854 under XGBoost. Under a third of a point, twice. The larger effect sits in who gets scored at all, the argument in the pipeline and model comparison. Neither study isolates individual attributes.
Education and employment attributes came out far stronger than expected. Di Maggio, Ratnadiwakara and Carmichael, in NBER Working Paper 29840 of March 2022, studied 770,523 personal loans and reported a test-sample AUC of 66.1 percent against 60.6 for a traditional model of roughly 1,400 credit-report variables, with education alone moving estimated default probability up to 4.2 percent. The alternative inputs are education, employment, loan purpose and device attributes, and no bank data.
The Consumer Financial Protection Bureau reported on 6 August 2019 that the same lender's model approved 27 percent more applicants at 16 percent lower average APRs, while stating that the simulations were not separately replicated by the Bureau and reflect the data and the method together.
Utility and telecom payment history separates borrowers hard, on behaviour recorded long ago. The Policy and Economic Research Council, a non-profit body, reported in May 2015 across roughly 4 million US credit files that consumers with a 90-day utility or telecom delinquency in the twelve months before July 2009 reached a 47.7 percent bank card delinquency rate, against 4.5 percent among consumers whose accounts were 24 months old or more with no delinquency ever reported. The window closes in June 2010, sixteen years ago, unreplicated at that scale.
Nine categories on one standard: measured against observed repayment on a lending population, isolated from other inputs, produced by someone with nothing to sell, and tested out of time. Vendor research is labelled in the cell.
| Category | What the published evidence shows | Strength on the four tests |
|---|---|---|
| Bank transaction (cash-flow) data | FinRegLab, July 2019, six non-bank lenders: standing alone, cash-flow metrics generally performed as well as traditional credit scores, among populations and products where traditional credit history is not available or reliable. FinRegLab, July 2025, out-of-time sample of 424,546 consumers: adding cash-flow to bureau data moved ROC-AUC 0.8653 to 0.8682 (logistic) and 0.8831 to 0.8854 (XGBoost) | Repayment yes, independent yes, out of time yes. Not isolated: attribute-level rankings withheld. Best-evidenced of the nine, and the incremental gain is under a third of a point |
| Education and employment attributes | Di Maggio, Ratnadiwakara and Carmichael, NBER WP 29840, March 2022, 770,523 loans: test-sample AUC 66.1% against 60.6% for a ~1,400-variable credit-report model; education alone moves estimated default probability up to 4.2%. CFPB, 6 August 2019: 27% more approvals at 16% lower APRs, not separately replicated by the Bureau | Repayment yes, independent yes, out of time not reported. Partly isolated: the alternative block is education, employment, loan purpose and device attributes. Closest thing here to a category-level measurement |
| Utility and telecom payment history | Policy and Economic Research Council, May 2015, about 4 million files: 47.7% bank card delinquency among consumers with a 90+ day utility or telecom delinquency in the 12 months before July 2009, against 4.5% among those whose accounts were 24+ months old with no delinquency ever reported | Repayment yes, independent yes. Not isolated: raw separation, not incremental lift. Observation window ends June 2010, sixteen years ago, non-peer-reviewed, no public replication |
| Digital footprint (device, email, channel) | Berg, Burg, Gombović and Puri, NBER WP 24551, April 2018, 254,808 German e-commerce observations from October 2015 to December 2016: AUC 68.3% bureau alone, 69.6% footprint alone, 73.6% combined | Repayment yes, independent yes, isolated yes. Out of time not reported. One market, one 14-month window, variables tied to 2015 and 2016 consumer technology |
| Psychometric questionnaires | Arraiz, Bruhn and Stucchi, World Bank Economic Review, 2016: among 1,993 Peruvian applicants, those the tool rejected but traditional scoring accepted were 8.6 points more likely to reach 90 days in arrears | Repayment yes, independent yes, isolated yes as a secondary screen. Out of time not reported. Small-business lending, one emerging market |
| Mobile phone metadata | World Bank Policy Research WP 9074, December 2019, 7,068 borrowers: AUC 0.760 on held-out folds within the sample period, 0.595 trained early and tested late, against a bureau baseline of 0.565 and 0.550 | The only category with a published out-of-time result, and it loses most of its edge. Repayment yes, independent yes, isolated yes. Bureau coverage in that market was weak |
| Rent payment history | Urban Institute randomised trial, 4 June 2025, 141 treatment and 128 control renters: share with no credit score fell 16% to 8% in treatment against 23% to 21% in control, and the near-prime share rose 40% to 57% against 38% to 40%, an estimated 12-point effect. The trial's estimated 7-point average score effect among renters who already had scores is not statistically significant. HUD analysis reported by Urban Institute, January 2025: FHA applicants accepted on rental history showed lower delinquency than the lowest-scoring accepted directly. TransUnion (vendor research), December 2021: 10+% improvement in predicting delinquencies over one year | Fails the repayment test. The trial measures whether a credit file forms, not whether it repays, and its score effect is not significant. The only lift figure found is vendor research |
| Neighbourhood and area aggregates | Hlongwane, Ramaboa and Mongwe, PLOS One, May 2024: a 22-variable alternative bundle including regional economic ratings and local population characteristics raised AUC from 0.779 to 0.794 on 356,255 customers | Repayment yes, but on a public competition dataset rather than a lender's book. Not isolated. Fails the proxy test by construction. Weakest of the nine |
| Social media content and graph | Ge, Feng, Gu and Zhang, Journal of Management Information Systems, 2017: self-disclosed social media activity predicted default on a peer-to-peer platform, with social stigma as the mechanism. Wei, Yildirim, Van den Bulte and Dellarocas, Marketing Science, 2016: borrowers who know their network is scored form fewer ties | Repayment yes, on a peer-to-peer platform. Fails the stability test by construction, since the signal changes once borrowers know it is in use |
Which signals only work in the conditions they were measured in?
Digital footprint variables, psychometric questionnaires and mobile phone metadata. All three isolate the category cleanly, which is more than the top three rows manage. All three were measured on one sample, in one market, in one period, and each carries a reason the result may not travel.
Digital footprint has the cleanest isolated measurement in the table. Berg, Burg, Gombović and Puri, in NBER Working Paper 24551 of April 2018, scored 254,808 German e-commerce observations from October 2015 to December 2016 at an AUC of 69.6 percent for the footprint alone, 68.3 for the bureau score, 73.6 combined. The variables carrying that load included device type and email provider, technology of 2015 and 2016.
Psychometric questionnaires came out better than expected. Arraiz, Bruhn and Stucchi, in the World Bank Economic Review in 2016, screened 1,993 Peruvian applicants and found that among entrepreneurs with credit history, those the tool rejected but traditional scoring accepted were 8.6 percentage points more likely to reach 90 days in arrears. Real incremental lift, from a questionnaire.
Mobile phone metadata is the only one of the nine with a published out-of-time result, and it fails it.
Which categories fail one of the four tests outright?
Rent payment history fails the lift test, neighbourhood aggregates fail the proxy test, and social media fails the stability test. Each has real published research behind it. In each case the research measures something other than what an underwriter needs it to measure, which is what one standard across nine categories is for.
Rent payment history has the cleanest experiment in the table and the wrong outcome variable. Theodos, Teles and Lieberman at the Urban Institute, in a randomised trial of 141 treatment and 128 control renters published 4 June 2025, found the share with no credit score fell from 16 to 8 percent in treatment against 23 to 21 percent in control, and the near-prime share rose from 40 to 57 percent against 38 to 40 percent. The same trial's estimated 7-point average score effect, among renters who already had scores, is not statistically significant, and the authors say so. Credit file formation matters, and it is not repayment. The HUD finding reported by Urban Institute on 16 January 2025, that FHA applicants accepted on rental history showed lower delinquency than the lowest-scoring accepted directly, reaches us secondhand. The much-quoted more than 10 percent gain is TransUnion vendor research.
Neighbourhood and area aggregates fail the proxy test by construction, because an area variable measures the area. Hlongwane, Ramaboa and Mongwe, in PLOS One in May 2024, raised AUC from 0.779 to 0.794 on 356,255 customers from a public competition dataset by adding 22 predictors including regional economic ratings. A bundle, on a dataset nobody lent against.
Social media has the strangest evidence base of the nine. Ge, Feng, Gu and Zhang, in the Journal of Management Information Systems in 2017, found self-disclosed social media activity predicted default on a peer-to-peer platform, with social stigma as the mechanism. Wei, Yildirim, Van den Bulte and Dellarocas, in Marketing Science in 2016, modelled borrowers who know their network is scored and found they form fewer ties. A signal that works only while borrowers do not know it degrades on disclosure.
What questions does a good data vendor answer easily?
Five. What population was the lift measured on, and did those people borrow? What was the out-of-time result rather than the cross-validated one? When was the model last refit? Which variables carry the weight? And what does the signal know about the applicant that the file does not already contain?
A category whose only lift measurement was produced by the firm selling it is unmeasured until somebody else measures it. That standard demoted rent payment history and promoted education and employment. A scorecard that only confirms the prior is not a scorecard.
What we couldn't verify
Three gaps sit under this table. No Canadian study measures predictive lift from any of these nine categories on a Canadian lending population, so every effect size above is American, German, Peruvian, Ethiopian or drawn from a public competition dataset.
No study isolates bank transaction data from every other input, leaving the best-evidenced category among the least precisely measured, and no independent lift measurement exists for rental trade lines.
Out-of-time results are missing almost everywhere, including for the signals inside a transaction history that a cash-flow underwriting policy rests on.
Common questions
What is alternative data in underwriting?
Alternative data in underwriting is any input about an applicant that does not come from a credit bureau file. The Federal Reserve Bank of Kansas City, in a June 2023 briefing by Terri Bradford, sorts it into financial data such as rent and transaction records, and non-financial data such as education.
Which alternative data has the strongest evidence behind it?
Bank transaction data, on two FinRegLab studies. The July 2019 work found cash-flow metrics standing alone generally performed as well as traditional credit scores for thin-file populations. Its July 2025 out-of-time study found adding cash-flow to bureau data moved ROC-AUC from 0.8653 to 0.8682, under a third of a point.
Does rent payment history predict credit performance?
No independent study measures that directly. The Urban Institute's June 2025 randomised trial measures file formation, with the unscored share falling 16 to 8 percent in treatment against 23 to 21 percent in control, and its average score effect is not statistically significant. The only lift figure is TransUnion vendor research.
Do education and employment data improve credit models?
Di Maggio, Ratnadiwakara and Carmichael, in NBER Working Paper 29840 of March 2022, reported a test-sample AUC of 66.1 percent against 60.6 percent for a roughly 1,400-variable credit-report model on 770,523 loans, using education, employment, loan purpose and device attributes and no bank data.
How can a lender tell whether an alternative data category will keep working?
By asking for the out-of-time result rather than the cross-validated one. Björkegren and Grissen, in World Bank Policy Research Working Paper 9074 of December 2019, reported a mobile-metadata AUC of 0.760 on held-out folds within the sample period and 0.595 when the model was trained early and tested late.
Carousel connects the bank data layer this evidence keeps pointing at. See how verification fits your flow


