How to Tune Transaction-Monitoring Scenarios and Cut AML False Positives

Every AML team knows the pain: a transaction-monitoring system that fires thousands of alerts, most of which turn out to be nothing. Left unchecked, false positives bury analysts, delay genuine investigations and inflate cost — without improving detection. The good news is that tuning is a structured, repeatable discipline, not guesswork. These are the lessons I’ve learned doing it in practice, written up so your team can apply them.

Why false positives pile up

Most monitoring scenarios ship with conservative, one-size-fits-all thresholds. They ignore customer segmentation, expected behaviour and the institution’s actual risk appetite. The result is alerts on perfectly normal activity — salary payments flagged as “structuring”, routine remittances flagged as “high-velocity”. Three patterns show up in almost every noisy system:

  • Global thresholds — one limit applied to a student account and a trading company alike. Neither is monitored well.
  • Scenarios nobody retires — rules added after every audit finding or typology paper, never reviewed again, overlapping and re-alerting on the same behaviour.
  • Data quality gaps — missing occupation codes, stale expected-turnover figures and unmapped transaction types force scenarios to fire “just in case”.

A tuning workflow that holds up

  1. Baseline — measure each scenario’s alert volume and true-positive rate over the last 6–12 months. You cannot improve what you haven’t measured.
  2. Segment — group customers by risk band, product and expected behaviour. Thresholds belong to segments, not to the whole bank.
  3. Hypothesise — propose a specific change (“raise threshold X for retail salary accounts”) with the expected effect on volume and detection.
  4. Test on history — replay the proposed setting against historical data and compare outcomes before touching production.
  5. Deploy and monitor — apply the change under change-control, then watch volumes and dispositions for a full cycle to confirm the effect matched the hypothesis.

Tune with evidence, not instinct

  • Segment first. Tune thresholds per customer risk band and product, not globally. The same scenario can be quiet and accurate in one segment and useless in another.
  • Use historical alert outcomes. Pull 6–12 months of alerts and their dispositions to see where productivity (true positives) actually lives — it is rarely where the noise is loudest.
  • Change one variable at a time so you can attribute the effect. Simultaneous changes make results unreadable and un-defendable.

Above-the-line and below-the-line testing

Before you move a threshold, sample alerts just above and just below the proposed line. Above-the-line testing confirms you still catch the risk you care about; below-the-line testing proves you are not about to miss productive alerts. Size the samples so the conclusion is statistically defensible, disposition them with your normal investigation standards, and record the results. This evidence is exactly what regulators and auditors expect to see behind every threshold change.

Data quality is half the battle

A surprising share of “tuning” problems are really data problems. If customer segments are stale, occupation and expected-activity fields are empty, or transaction codes are mapped inconsistently, no threshold will behave sensibly. Fix the feeds — and the KYC data they depend on — alongside the scenarios, or you’ll tune yourself into a corner. The same discipline applies to sanctions and PEP screening, where matching quality lives and dies on identifiers.

Make it audit-ready

Document the rationale, the data sample, the testing and the sign-off for every change in a scenario register. When the regulator or internal audit asks “why is this threshold set here?”, the answer should already be written down. A good register carries, at minimum:

  • Scenario purpose — the typology or risk it exists to detect, in one sentence.
  • Current parameters — thresholds, segments and the date they took effect.
  • Evidence — the analysis and above/below-the-line results behind each change.
  • Approvals — who proposed, who reviewed, who signed off, under which governance forum.
  • Review date — when this scenario will next be re-validated.

This is the same “understand your risks, manage them deliberately, prove it” discipline that underpins an ISO 27001 ISMS — good governance looks the same whichever control you point it at.

Common mistakes to avoid

  • Tuning to reduce volume alone. The goal is better detection per analyst hour, not a smaller queue at any cost.
  • Skipping below-the-line testing because “the change is obviously safe”. That sentence appears in a lot of audit findings.
  • One-off projects. Behaviour, products and typologies drift; tuning is an annual-at-least cycle, not a rescue mission.

These notes are educational — the field lessons of a practitioner, shared for teams building the same muscles. If you’re a fintech starting from zero, begin with AML compliance for fintechs and startups. And if your challenge is on the cybersecurity, GRC or data-protection side, that’s exactly what I offer — see my services or get in touch.

Bader Alkandery

Freelance cybersecurity, GRC & data-protection consultant in Kuwait — MSc Cyber Security & Networks (Best Paper), CompTIA Security+.

Keep reading

More insights