Skip to main content
Category: Risk Assessment

Below-the-Line Testing

Also known as: BTL, Below-the-line test, BTL testing
Simply put

Below-the-line testing is a check that AML teams run on transactions and customer activity that fall just under the thresholds their monitoring rules use to flag suspicious behavior. By lowering those thresholds and sampling the activity that would normally pass unnoticed, teams can see whether genuinely suspicious behavior is slipping through undetected. It is typically used to help tune, or calibrate, monitoring rules so they capture the right activity.

Formal definition

Below-the-line (BTL) testing is a rule-calibration and validation technique used within AML transaction monitoring to examine activity that falls below the current alerting thresholds or criteria and would therefore not generate an alert under existing rule parameters. Practitioners generally lower or relax thresholds below the established baseline and sample the additional population captured, assessing whether that population contains genuinely suspicious activity that the current settings fail to detect and, correspondingly, identifying the point at which thresholds may be set. BTL testing is commonly paired with above-the-line (ATL) testing, which examines the effect of raising thresholds, as part of an overall calibration process for rule-based monitoring systems. It is an operational optimization and effectiveness-testing measure intended to help detect and mitigate the risk of missed suspicious activity; it is a tuning exercise rather than a legal test, and a sampled transaction falling below the line does not itself indicate wrongdoing. Specific thresholds, sampling methodologies, and calibration frequency vary by institution, monitoring system, and applicable supervisory expectations, and should be confirmed against the relevant regulatory and internal-model-governance requirements.

Why it matters

Rule-based transaction monitoring systems rely on thresholds and criteria to decide which activity warrants an alert. Any threshold, however well designed, draws a line: activity that sits just beneath it passes without generating an alert. Below-the-line (BTL) testing addresses the risk hidden in that gap by examining the population of transactions that fall below current alerting parameters, helping institutions understand whether genuinely suspicious behavior is escaping detection simply because it does not cross a configured threshold. Without such testing, a monitoring program may appear to be functioning while systematically missing structured or deliberately calibrated activity designed to stay under the line.

BTL testing is closely tied to the effectiveness and defensibility of an AML program. Supervisors generally expect obliged entities to be able to demonstrate that their monitoring rules are calibrated on a reasoned basis rather than left at default or arbitrary settings. By sampling the additional activity captured when thresholds are lowered, teams gather evidence about where the threshold should sit and can document the rationale behind their calibration decisions. Paired with above-the-line (ATL) testing, which examines the effect of raising thresholds, BTL testing forms part of an overall calibration and validation process that supports model governance and tuning.

It is important to keep BTL testing in its proper frame: it is an operational optimization and effectiveness-testing measure intended to help detect and mitigate the risk of missed suspicious activity, not a legal test. A transaction that falls below the line and is surfaced during sampling does not by itself indicate wrongdoing; it is simply activity that the current settings did not flag, examined so the institution can decide whether its thresholds are appropriate. Specific thresholds, sampling methodologies, and calibration frequency vary by institution, monitoring system, and applicable supervisory expectations, and should be confirmed against the relevant regulatory and internal model-governance requirements.

Who it's relevant to

Transaction Monitoring and Tuning Teams
Analysts and specialists responsible for calibrating rule-based monitoring systems use BTL testing to examine activity below current thresholds and to determine where thresholds may appropriately be set. It is a core part of their tuning workflow, typically paired with above-the-line testing.
Model Validation and Governance Functions
Those charged with validating and governing monitoring models rely on BTL testing as evidence that thresholds are set on a reasoned basis rather than left at defaults, supporting documentation of calibration decisions in line with applicable model-governance requirements.
AML Compliance Officers and MLROs
Compliance leaders responsible for the effectiveness of the monitoring program use the results of BTL testing to demonstrate that the institution actively assesses the risk of missed suspicious activity and calibrates its rules accordingly. They should note that a below-the-line sample does not itself indicate wrongdoing.
Internal Audit and Independent Testing
Functions performing independent review of the AML program may examine whether BTL testing is conducted, how sampling is performed, and how frequently calibration occurs, confirming these against internal policy and supervisory expectations.

Inside BTL

Below-the-Line Threshold
The point beneath a transaction monitoring system's alerting threshold at which activity does not generate an alert. Below-the-line testing focuses on this population of transactions or scenarios that fall short of the configured parameters and are therefore not surfaced for review under normal operation.
Threshold and Parameter Tuning
The practice of assessing whether monitoring rule thresholds, values, and parameters are appropriately calibrated. Below-the-line testing samples activity just under those thresholds to determine whether potentially suspicious behaviour is being systematically missed, informing whether parameters should be lowered or otherwise adjusted.
Sampling Methodology
The approach used to select non-alerting transactions for review, which may be statistical, risk-based, or a combination. The methodology defines the sample size, the segments examined, and the period covered, and is typically documented so results can be reproduced and defended.
Effectiveness Evaluation
Analysis of whether the sampled below-the-line activity contains items that, on review, would warrant an alert or further investigation. This helps evaluate how effectively the monitoring configuration detects potentially suspicious activity, rather than confirming that any single transaction is unlawful.
Model Governance Linkage
The connection between below-the-line testing and broader model validation and governance frameworks. Findings generally feed into documentation, change management, and periodic review of the transaction monitoring model, supporting a risk-based approach to calibration.
Contrast with Above-the-Line Testing
Above-the-line testing examines activity that did generate alerts to assess whether thresholds are set too low or produce excessive noise, while below-the-line testing examines non-alerting activity to assess whether thresholds are set too high and miss relevant behaviour. The two are complementary but distinct exercises.

Common questions

Answers to the questions practitioners most commonly ask about BTL.

Does a below-the-line alert that a customer's transaction generated mean the customer was engaged in suspicious activity?
No. Below-the-line testing examines transactions or behaviors that fell just under a monitoring threshold and therefore did not generate a productive alert. Identifying that activity sat below the line is an assessment of whether the threshold is calibrated appropriately; it is not a determination that any customer engaged in wrongdoing. A hit surfaced during such testing may warrant further review, but on its own it does not establish suspicion, let alone criminal conduct. Any escalation would follow the institution's normal investigation and, where applicable, suspicious activity or suspicious transaction reporting processes.
Is below-the-line testing the same thing as tuning or validating a transaction monitoring system?
Not exactly, though they are related. Below-the-line testing is one technique used within the broader activities of monitoring optimization and model validation. It specifically focuses on the population of activity that fell below existing alerting thresholds to assess whether those thresholds may be set too high and are missing potentially relevant activity. Above-the-line testing, sampling of alerted activity, scenario coverage assessments, and data-quality checks are separate components. Treating below-the-line testing as synonymous with the full tuning or validation exercise can leave other calibration and coverage questions unaddressed.
How is a below-the-line sample typically selected?
Approaches vary by institution and are generally shaped by the entity's risk-based methodology rather than a single prescribed formula. Common practice is to define a band of activity sitting beneath the current threshold and draw a sample from that population, often using statistically informed sampling to support conclusions about the wider set. The design typically considers the relevant scenario or rule, the customer or product segment, and the size of the below-the-line population. Because methodologies differ and expectations should be confirmed against applicable supervisory guidance, the sampling rationale is usually documented so it can withstand review.
What results from below-the-line testing might indicate that a threshold should be lowered?
Where testing surfaces activity that appears consistent with the risks a scenario is designed to detect but which fell below the current threshold, that can suggest the threshold may be set too high and warrant recalibration. The judgment is qualitative and risk-based rather than mechanical: a small number of items requiring review does not automatically justify a change, and any adjustment is generally weighed against alert volume, investigative capacity, and the institution's risk appetite. Conclusions are typically supported by documented analysis rather than a single fixed pass or fail figure.
How often should below-the-line testing be performed?
There is no universally mandated frequency. Institutions generally set a cadence within their model risk management or monitoring governance framework, and testing is often triggered periodically as well as by events such as material changes to products, customer base, typologies, data feeds, or the monitoring system itself. The appropriate frequency is typically informed by the institution's risk profile and any applicable supervisory expectations, which should be confirmed against the relevant regime rather than assumed to be a standard interval.
How should below-the-line testing be documented and governed?
Documentation generally covers the objective, the definition of the below-the-line population, the sampling methodology and rationale, the results, the disposition of any items reviewed, and any resulting threshold changes with the supporting justification. Governance typically involves oversight by relevant functions such as compliance and independent validation or audit, with results reported through the institution's model risk or AML governance channels. Maintaining a clear, reviewable record supports the ability to demonstrate to auditors and supervisors that monitoring thresholds have been assessed as part of a risk-based program, though it does not by itself guarantee that all relevant activity is detected.

Common misconceptions

Below-the-line testing that identifies missed activity means the institution has failed to file required reports and has committed a violation.
Identifying transactions that arguably should have alerted is an input to calibration and effectiveness review, not proof of wrongdoing or a filing failure. Whether any activity is genuinely suspicious, and whether a report is warranted, remains a separate determination made through investigation and professional judgement.
Below-the-line testing produces a definitive, exhaustive measure of everything the monitoring system missed.
It is generally a sampling-based exercise covering a subset of non-alerting activity over a defined period. It provides evidence about calibration and detection effectiveness for the tested population but does not exhaustively capture all potentially missed activity, and its conclusions depend on the sampling methodology chosen.
Below-the-line testing is universally mandated in identical form across all jurisdictions and obliged entities.
It is primarily a risk management and model governance practice rather than a single prescriptive legal requirement. Expectations around transaction monitoring tuning and testing vary by regime and by the nature and size of the obliged entity, and specific supervisory expectations should be confirmed against the applicable rules and guidance.

Best practices

Document the sampling methodology, including sample size, segmentation, and the review period, so that below-the-line testing can be reproduced and defended to auditors and supervisors.
Pair below-the-line testing with above-the-line testing so that both under-alerting (missed activity) and over-alerting (excessive noise) are assessed as part of holistic threshold tuning.
Feed testing results into the institution's model validation and change management processes so that any parameter adjustments are governed, evidenced, and periodically reviewed.
Frame findings as calibration and effectiveness evidence rather than as determinations of suspicion, and route any genuinely concerning items into the standard investigation and escalation workflow.
Apply a risk-based lens when selecting segments and thresholds to test, prioritising higher-risk products, customers, and channels where missed activity would be most consequential.
Retain clear records of rationale where thresholds are left unchanged after testing, so the decision demonstrates a considered, risk-based approach rather than an omission.