The Doctored Paystub Economy: Inside Document Tampering at Scale
What tamper detection inspects when a doctored paystub arrives, why AI-edited files now beat it, and where document checks stop.
Alfred BEditorial Reviews
A lending file is a stack of documents the applicant chose to send. Underwriting, pricing, funding and the audit trail all sit on top of that stack, and every one of them inherits the same assumption. The paper is what it says it is.
A doctored paystub is an income document altered after issuance. Detection systems inspect four things: how the file was built, whether the page renders consistently, whether the document's own arithmetic holds, and how closely it matches the issuer's template. Verified source data removes the question rather than scoring it.
What follows is what the public evidence supports about the scale, what the inspection layers do, and where they stop.
How much of lending fraud is a document problem?
Most of Fannie Mae's confirmed mortgage fraud findings, on the only long-run public count available. Fannie Mae's Mortgage Fraud Investigative Findings 2005–2021 records income and employment misrepresentation appearing in 12% of its confirmed fraud findings in 2005 and 61% in 2020. No other misrepresentation category in that series grew at anything close to the same rate.
The largest single measurement of what happens when document review gets compressed comes from the pandemic loan programs. The US Small Business Administration's Office of Inspector General reported in June 2023, in Report 23-09, that over $200 billion in potentially fraudulent EIDL and PPP loans was disbursed, roughly 17% of about $1.2 trillion, and attributed the exposure to controls weakened to move money quickly.
The US Financial Crimes Enforcement Network issued an alert on 13 November 2024, FIN-2024-Alert004, describing criminals using generative AI to alter or generate images used for identification documents, and noting increased suspicious activity reporting on the pattern beginning in 2023. The detection route matters as much as the scheme. Institutions found these cases through inconsistencies between one submitted document and another, or between a document and the rest of the customer's profile, rather than by reading any single document more closely.
Canadian figures are thinner and mostly vendor-produced. In vendor research published on 15 April 2026, Equifax Canada reported that falsified financial information appeared in 21% of first-party banking and deposit fraud cases in the fourth quarter of 2025, against 1.5% a year earlier, while mortgage application fraud fell 12.5% year over year.
Police-reported crime counts a different universe from lending. Statistics Canada reported on 22 July 2026 that Canada's police-reported fraud rate fell 4% in 2025, to 492 incidents per 100,000 population, still 61% above 2015.
What does fake paystub detection actually inspect?
Four families of signal, none of them decisive on its own. Detection systems examine the file's internal construction, the consistency of how text is rendered across a page, whether the stated figures reconcile with each other and with adjacent periods, and whether the layout conforms to what the named employer or payroll processor produces.
The rendering family is the one computer vision research has studied most. Chenfan Qu and co-authors, in work presented at the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, built the DocTamper dataset of 170,000 document images and set out the underlying signal plainly: changing text inside a compressed document disturbs the distribution of the file's compression coefficients, leaving discontinuities between changed and unchanged regions. The same paper explains why documents are harder than photographs. A page has one background colour and one font cluster, leaving far less natural variation for a detector to work with.
The arithmetic family is the least interesting and the most durable. A pay statement is a small closed system of numbers that has to agree with itself and with the periods either side of it. That check needs no forensics, only the rest of the file.
Template conformance compares what arrived against the known output of the institution it claims to come from, so its coverage is a function of the reference library. A small employer running payroll on a spreadsheet has no template to match.
In production these run as layers and return a score rather than a verdict. That score is an input to a decision about whether to ask for something else, and the something else is usually another document.
The four inspection families used in fake paystub detection, and the question each one leaves unanswered.
| Layer | What it inspects | What it cannot settle |
|---|---|---|
| File construction | How the document was produced, and whether its internal structure is consistent with a single generation event | Whether a cleanly regenerated file reflects real earnings |
| Rendering consistency | Compression and glyph-level discontinuities across the page (Qu et al., CVPR 2023) | Files with no capture history to compare against |
| Internal arithmetic | Whether gross, deductions, net and year-to-date figures reconcile within and across periods | Whether consistent figures correspond to real earnings |
| Template conformance | Whether layout matches the known output of the named issuer | Any issuer absent from the reference library |
| Source verification | Nothing about the document; retrieves the underlying record from the institution holding it | Applicants who will not, or cannot, connect |
Why is automated detection getting harder?
Because the detectors were trained on a kind of tampering that is being replaced. AIForge-Doc, a benchmark posted to arXiv on 24 February 2026 by Jiaqi Wu and co-authors, tested two forgery detectors and one general multimodal model against AI-inpainted edits to financial and form documents. All three degraded, and one landed at roughly chance.
On the AIForge-Doc benchmark, the general-purpose forgery detector TruFor scored an AUC of 0.751, against 0.96 on the conventional NIST16 forgeries. DocTamper, trained specifically on documents, scored 0.563 on AIForge-Doc against 0.98 on its own test set. GPT-4o, a general model the AIForge-Doc authors prompted zero-shot and which was never trained on forgery detection, scored 0.509, essentially chance.
A separate group found part of the reason. Zeqin Yu and co-authors, in a NeurIPS 2025 datasets paper, assembled 294,182 tampered text images, including 16,750 real tampering instances produced by 67 human participants whose editing sessions were recorded. Models trained on synthetic tampering transferred poorly to the real thing. Retraining on the more realistic data lifted average F1 across methods by more than 14 points.
Put those results side by side and the pattern is not that detection fails. It is that detection dates. A detector encodes the manipulation techniques that existed when its training set was assembled, and the techniques keep moving.
A view rather than a finding: document inspection earns its keep and depreciates while it does so. It catches real cases cheaply, it is the only check available when there is nothing to connect to, and its maintenance bill comes due every year against methods nobody has seen yet.
What does verifying at source change?
It changes who the lender is trusting. A source-verified income record arrives from the payroll processor, the bank or the tax authority rather than from a file the applicant assembled. Fannie Mae prices that difference directly, offering representation and warranty enforcement relief on components validated through its Desktop Underwriter validation service.
Fannie Mae's Selling Guide section B3-2-02, in the version dated 5 February 2025, states that when income, employment or assets are validated through the DU validation service, the lender may receive representation and warranty enforcement relief related to that component, and that validated employment satisfies the verbal verification of employment requirement. The largest buyer of US mortgages assigns a different risk weight to the same fact depending on where the fact came from.
The US tax authority runs the same pattern. The Internal Revenue Service's Income Verification Express Service lets a taxpayer authorise a lender to receive return, W-2 and 1099 transcripts directly from the IRS, and the IRS states that it sends those records only if the taxpayer approves the request. Consent is the gate. No document sits in the path.
There is quieter evidence in how the largest counterparty behaves when documents fail. In a fraud alert dated 1 July 2021, Fannie Mae published 63 businesses given as borrowers' places of employment on loans originated between 2015 and 2019 whose existence it could not confirm, and told lenders to verify that the place of employment actually exists. The response to unverifiable paper was to go around the paper.
Where does Canada sit on this?
Behind, and openly working on it. The Canada Revenue Agency ran a mortgage industry consultation on a potential income verification tool and published its findings in July 2025. As of August 2026 there is no Canadian equivalent of a tax authority delivering income transcripts to a lender on the taxpayer's authorisation.
The CRA's "What we learned" report records 1,637 completed questionnaire submissions, 61% of them from mortgage brokers and 51% from Ontario. Participants described fake or altered documents being used to inflate borrower income, and identified Canadian tax documents as central to reducing mortgage application fraud. On access, 85% supported banks using such a tool and 75% supported mortgage brokers.
Bank data in Canada is further along than payroll or tax data. Draft Consumer-Driven Banking Regulations were published in the Canada Gazette, Part I on 27 June 2026, covering profile, account, balance, transaction and product data from deposit, payment, investment and lending accounts, with the first phase limited to read-only access. Transaction data shows the deposit that a pay statement claims to explain, which for many files is the more useful of the two records.
What we couldn't verify
No public source establishes what share of income documents submitted to lenders has been altered. Every figure we found on that question came from a vendor reporting its own detection rate on its own book, which cannot be checked from outside and is not comparable between vendors. If someone quotes a percentage of fake paystubs in circulation, the number almost certainly has no reachable origin.
The 2021 column of the Fannie Mae findings table did not render consistently across the copies we retrieved, so the figures above stop at 2020.
The academic benchmarks measure receipts and forms rather than pay statements. We found no published detection-accuracy figure for paystubs specifically, so the AIForge-Doc and DocTamper results are indicative here rather than direct.
And nothing published lets a Canadian lender compare defect rates on source-verified files against document-verified ones. The gap in the evidence sits on measurement rather than on whether tampering happens at all.
Common questions
What is fake paystub detection?
Fake paystub detection is the practice of testing a submitted income document for evidence it was altered after it was issued. Systems inspect file construction, rendering consistency across the page, whether the stated figures reconcile, and whether the layout matches the named issuer's known output.
Can software reliably detect an altered paystub?
Not against current methods. The AIForge-Doc benchmark, published in February 2026, found the document-trained detector DocTamper scoring an AUC of 0.563 on AI-edited documents against 0.98 on its own test set, and the general model GPT-4o, prompted zero-shot, scoring 0.509.
What do tamper-detection systems look at?
Four families of signal: how the file was constructed, compression and glyph-level inconsistency across the page, whether gross, deductions, net and year-to-date figures reconcile, and conformity to the issuer's template. Each returns a probability. None of them is decisive alone.
Is there Canadian data on document fraud in lending?
Very little. The Canada Revenue Agency's July 2025 consultation report on a potential income verification tool records industry accounts of fake or altered income documents across 1,637 questionnaire responses. Beyond that, most Canadian figures on the subject are vendor-published, including Equifax Canada's April 2026 fraud research.
What does source verification replace?
It replaces the document with the record behind it. Fannie Mae's Selling Guide, dated 5 February 2025, offers representation and warranty enforcement relief on income, employment and asset components validated through its Desktop Underwriter validation service, which is the same fact carrying a different risk weight.
Carousel collects income and bank data through source connections and document capture in the same intake flow. See how verification fits your flow


