Dolores "Lola" Umbral spent eighteen years as a BSA/AML examiner before crossing into independent consulting. Her career has been defined by a single preoccupation: what happens when you try to assess a system's adequacy using only the evidence that system chose to produce. We met her at a café where she immediately ordered two coffees, explaining that threshold questions require double the caffeine.
A note on Lola: she exists as a composite drawn from publicly documented examination practices, enforcement actions, and the FFIEC examination manual. Her observations are real. Her person is assembled.
You've described your former job as "auditing a flashlight by looking at what it illuminates." What do you mean?
Lola: When I walk into a bank to examine their AML program, what do they hand me? Alert logs. Case files. SAR filings. Disposition records. Every piece of that was generated by a monitoring system that someone else configured, with thresholds someone else set, running against data feeds someone else validated. Or didn't validate.
So I'm looking at the output of a machine and being asked to judge whether the machine is working. Using the machine's own output.
You see the problem.
That sounds like a structural impossibility.
Lola: It would be, if I only looked at what they gave me. The trick, and it took me years to learn this, is that what the system didn't flag tells you more than what it did. We call it below-the-line testing.1 You pull a sample of transactions that scored just under the alert threshold. If you find suspicious activity down there, the threshold is wrong. And if the threshold is wrong, then every clean case file the institution ever showed an examiner was accurate about the cases it contained. It just didn't contain the right cases.
How often does below-the-line testing actually find problems?
Lola: More often than anyone's comfortable admitting. I can't give you a number, and that's the whole point, right? False negatives are structurally hidden.2 You only know the system missed something when someone outside the system finds it. Law enforcement calls with a lead. An examiner like me pulls the raw transaction data and spots clustering. Or, and this is the ugly one, a bank gets fined years later and everyone discovers the monitoring was blind the entire time.
Commonwealth Bank of Australia. Seven hundred million dollars.3 NatWest. Two hundred sixty-five million pounds flowing through and nobody's system blinked.4 Those institutions had case files. They had SARs. They had documentation. The monitoring system was running, alerts were being reviewed, cases were being closed with rationale. Everything looked fine from inside the flashlight's beam.
So a perfectly documented compliance program can coexist with massive failure?
Lola: Absolutely. And people act surprised every time, which is the part that gets me. The investigation record proves investigation quality. It does not prove population completeness. Those are different claims, and people confuse them constantly.
You mentioned pulling raw transaction data. Walk me through what you're actually looking for.
Lola: The FFIEC manual gives examiners authority to request complete transaction reports, not just alert outputs.5 I want the full funds transfer record: names, countries, amounts, dates, account types. And I want it sorted to show me where information is missing. Gaps in originator or beneficiary details, for instance. That might look like sloppy recordkeeping, but it can also mean the monitoring system isn't even ingesting those transactions. They're invisible before they're incomplete.
But here's where it gets interesting. I'm also looking at distribution shapes. You know what structuring looks like in the data? A cluster of transactions at $9,500, $9,800, $9,900. Just below the $10,000 reporting threshold.6 No single transaction triggers anything. But the pattern is screaming. The system can't hear it because each transaction is evaluated individually against a threshold, and each one passes. The criminal has learned the system's vocabulary and is speaking just below it.
The system teaches people how to evade it.
Lola: Congress literally had to criminalize structuring in 1986 because the reporting threshold created its own evasion technique.7 The threshold was the lesson plan. Which tells you something about the relationship between a rule and the behavior it's supposed to catch. Set a bright line, and people learn to stand just on the other side of it.
What's the most counterintuitive thing you've learned doing this work?
Lola: That dismissals matter more than filings.
Everyone focuses on SARs. Did you file, was the narrative complete, was it timely.8 Important questions. But I learned to spend most of my time on the cases that were closed without escalation.
Because a dismissal is a claim that nothing suspicious was found, and that claim is only as good as the evidence reviewed, which is only as good as the data the system surfaced, which is only as good as the threshold calibration, which, in a lot of institutions, has never been independently validated.
I once... well, I can't tell you specifics. But I've seen institutions where the documented rationale for a threshold was essentially "industry standard." That's not a rationale. That's a shrug.9
Does investigation work ever feed back into improving the monitoring system?
Lola: [long pause]
No.
That's the thing that still bothers me. There's typically no feedback loop from investigations back into the monitoring system.10 What I learn from reviewing a case, the patterns I notice, the data gaps I identify, the near-misses, none of that automatically adjusts what the system shows the next analyst. The system learned its calibration from some historical moment. Investigations happen in a different moment. But the threshold sits there, unchanged, like a fossil.
Some institutions do periodic tuning. The good ones test iteratively: lower the threshold, review what surfaces, lower it again until they stop finding suspicious activity below the line.1 But plenty of them are caught in a cycle. Tune, discover the old threshold was wrong, do a look-back of everything that was missed, retune, discover that threshold was wrong...
What's a look-back like?
Lola: Expensive. Brutal. You're reconstructing investigations that should have happened months or years ago, against evidence that may have degraded, for customers who may have moved on. You're building cases backward through time, against a monitoring system that, for the entire period in question, was producing clean records. The institution's own documentation says everything was fine. You're proving it wasn't.
Does that ever feel futile?
Lola: It feels necessary. Which is different.
If you could change one thing about how monitoring systems work, what would it be?
Lola: I'd make the system report on what it chose not to show me.
Right now I get a list of alerts. I want a list of near-misses. I want the system to say: "Here are the 4,000 transactions that scored between 70 and 80 percent of the alert threshold. Here's their distribution by customer type, geography, product."
Give me the shape of the shadow, not just the light.
Because my expertise, and I've spent two decades building it, operates inside a boundary I can't see from inside. The only way to find the boundary is to step outside the system's output entirely. And most people never think to do that. Why would they? The case files look complete. The documentation is thorough. Everything below the line is invisible by design.
Lola finished her second coffee and noted, with characteristic precision, that threshold-based beverage consumption, exactly two cups and never three, was itself a form of structured activity that no monitoring system would flag. She paused. "Though if I ordered two coffees at nine different cafés in the same morning, someone should probably look into it."
Footnotes
-
ACAMS, "Auditing the AML/CTF Transaction Monitoring System," https://www.acams.org/sites/default/files/2020-07/white-paper-Jon-Harvey.pdf ↩ ↩2
-
WorkFusion, "Overcoming AML Transaction Monitoring Challenges with AI," https://www.workfusion.com/blog/overcoming-aml-transaction-monitoring-challenges-with-ai/ ↩
-
Ankura, "When the Model Misses: Why AML Validation Can No Longer Be an Afterthought," https://ankura.com/insights/when-the-model-misses-why-aml-validation-can-no-longer-be-an-afterthought ↩
-
Chainalysis, "What Is Transaction Monitoring? AML Compliance, KYT & Crypto Monitoring Guide," https://www.chainalysis.com/glossary/transaction-monitoring/ ↩
-
FFIEC BSA/AML Appendix O — Examiner Tools for Transaction Testing, https://bsaaml.ffiec.gov/manual/Appendices/16 ↩
-
Arxiv, "Searching for Smurfs: Testing if Money Launderers Know Alert Thresholds," https://arxiv.org/pdf/2309.12704 ↩
-
Protiviti, "Tuning Suspicious Transaction Monitoring Scenarios," https://www.protiviti.com/sg-en/whitepaper/tuning-suspicious-transaction-monitoring-scenarios-combining-aml-expertise-and-data ↩
-
FFIEC BSA/AML — Suspicious Activity Reporting, https://bsaaml.ffiec.gov/manual/AssessingComplianceWithBSARegulatoryRequirements/04_ep ↩
-
Deloitte Switzerland, "AML Transaction Monitoring: Calibration of rule-based Transaction Monitoring vendor systems," https://www.deloitte.com/ch/en/Industries/financial-services/blogs/calibration-of-rule-based-transaction-monitoring-vendor-systems.html ↩
-
Stout, "Modernizing AML Transaction Monitoring Model Risk Management," https://www.stout.com/en/insights/article/modernizing-aml-transaction-monitoring-model-risk-management ↩
