Independent editorial research

AML Software Evaluation Framework

A structured method for turning financial-crime risks into testable software requirements, evidence requests, proof-of-concept measures, and an accountable buying decision.

By AML Tech Reviews Editorial TeamPublished Updated

A procurement team can receive three polished demonstrations and still be unable to answer one basic question: which control will work on the institution’s data, risks, and operating model? This framework turns that question into a documented evaluation. It is for compliance, technology, operations, procurement, internal audit, and model-risk stakeholders who need a defensible decision rather than a feature comparison.

Software is only one part of an AML control. Policies, data, investigators, escalation paths, quality assurance, and governance determine whether a system produces useful outcomes. The FATF Recommendations set a risk-based international standard, but they do not prescribe one product architecture. The evaluation should therefore begin with the institution’s obligations and assessed risks, not a vendor’s module list.

Decision context

Write a one-page decision statement before issuing a request for information. It should name the control being changed, the weakness or growth constraint behind the change, the affected business lines, and the evidence required for approval. A usable statement might say that the institution needs to replace a monitoring platform because material data fields cannot be traced from source to alert, scenario changes take too long to test, and investigation queues exceed approved service levels.

Separate four possible decisions:

  • retain and tune the existing system;
  • add a specialist component while keeping the core platform;
  • replace one control, such as screening or monitoring;
  • replace an integrated suite of controls and workflow tools.

These options carry different migration, validation, and concentration risks. An attractive new interface is not evidence that a full-suite replacement is justified.

Record the jurisdictions, regulated entities, products, customer types, channels, transaction volumes, languages, and data-residency constraints in scope. Map each requirement to a named risk or operational need. The FCA’s Financial Crime Guide asks firms to understand automated monitoring capabilities and limitations, keep customer information current, and understand rule and threshold rationales. Those expectations translate into evaluation questions about traceability, change control, and evidence—not a mandate for a particular technical design.

Requirements checklist

Use requirements that can be observed or tested. Mark each one mandatory, desirable, or out of scope.

Control coverage

  • Define the events the system must evaluate: onboarding, customer changes, list updates, transactions, payment messages, periodic reviews, or investigator feedback.
  • Specify the risks, typologies, jurisdictions, and business lines covered by each control.
  • State when a decision must be synchronous and when a queued process is acceptable.
  • Require a documented path from detected activity to disposition, escalation, reporting, and control feedback.

Data and integration

  • Inventory required source fields, owners, formats, frequencies, quality checks, and permitted uses.
  • Require field-level lineage from source through transformation to decision output.
  • Define behavior for late, missing, duplicated, malformed, or corrected records.
  • Specify reconciliation, replay, exception handling, and monitoring of failed feeds.
  • Identify authoritative systems for customer, account, transaction, identity, and reference data.

Configuration and governance

  • Require versioning, approval, testing, release, rollback, and effective dates for configuration changes.
  • Separate maker, checker, administrator, investigator, auditor, and read-only permissions.
  • Preserve the input, configuration version, output, user action, rationale, and timestamps needed to reconstruct a decision.
  • Define retention and export needs according to applicable law and institutional policy.

Operations and resilience

  • Set measurable availability, recovery, latency, throughput, queue, and support requirements.
  • Define priority handling without allowing priority to hide aging work.
  • Require management information for control coverage, data failures, alert volumes, dispositions, aging, overrides, and quality findings.
  • Test accessibility, usability, and investigator navigation with real job roles.

Evidence to request

Ask for evidence before accepting a claim. A response marked “supported” should point to a document, configuration screen, test result, or contractual term.

Request architecture and data-flow diagrams; a logical data model; API and batch specifications; permission and audit models; release notes; supported deployment patterns; incident and continuity procedures; and a sample service-level agreement. For hosted services, request current independent assurance reports and the supplier’s control-responsibility matrix. Review scope, exceptions, and customer responsibilities rather than treating a certification logo as a conclusion.

For analytics, request model or rule documentation, intended use, training and validation approach where applicable, performance measures, segmentation, known limitations, change history, and monitoring procedures. The Wolfsberg statement on monitoring effectiveness emphasizes outcomes, false-negative analysis, contextual information, and clear documentation for machine-learning models. It is industry guidance, not law, but it provides useful evidence prompts.

Ask for two customer references with comparable scale or complexity only when reference contact is permitted. Treat a reference as experience, not proof that the same implementation will work in a different data environment.

Proof-of-concept test plan

Do not let the supplier choose only the test data or success measures. Build a controlled dataset that represents normal activity, known risk examples, data-quality failures, boundary conditions, and operational volume. Mask or synthesize data where required, then verify that the transformation preserves the properties being tested.

Run the proof of concept in six stages:

  1. Data acceptance. Load representative records and reconcile counts, required fields, duplicates, rejected records, and transformations.
  2. Control behavior. Exercise mandatory scenarios, matching cases, risk changes, and configured workflows. Include expected non-alerts as well as expected alerts.
  3. Explanation. Ask an investigator to reconstruct why each sampled output occurred, which data contributed, and which configuration version applied.
  4. Change control. Modify a rule, threshold, workflow, or model setting through the proposed approval path. Test impact analysis, deployment, and rollback.
  5. Operations. Measure latency, throughput, queue behavior, navigation time, handoffs, and evidence capture under realistic load.
  6. Failure and recovery. Interrupt a feed, submit malformed data, retry a job, restore service, and reconcile what was or was not processed.

Define measures before testing. Useful measures include data acceptance and reconciliation rates, detection results on labelled cases, false-alert rate on the test set, time to explain a decision, investigator handling time, configuration deployment time, and recovery results. A test-set result is not a general performance guarantee. Document sample composition and uncertainty beside every measure.

Commercial and implementation questions

Pricing must be modeled against expected growth, not only current volume. Ask what is charged: records, transactions, API calls, screened names, users, environments, modules, data lists, storage, implementation days, or support tiers. Identify minimum commitments, overage treatment, indexation, currency, taxes, and renewal terms. Obtain a three-to-five-year cost model with low, expected, and high volume cases.

Implementation questions should expose work that falls between parties:

  • Who maps, cleans, and reconciles source data?
  • Who writes and validates control logic?
  • Which configurations are reusable across entities and jurisdictions?
  • What environments, test tools, and migration utilities are included?
  • How are open alerts and historical decisions migrated or retained?
  • What training is role-specific, and what documentation remains after go-live?
  • Which subcontractors handle data or provide critical services?
  • What assistance is available for exit, data export, and transition?

FinCEN’s culture-of-compliance advisory links effective programs to leadership, information sharing, adequate human and technological resources, and independent testing. That makes staffing, ownership, and testing capacity part of the commercial decision, not post-contract details.

Red flags

Pause the evaluation when a supplier will not define a claimed performance measure, disclose the test population, demonstrate audit history, or identify customer responsibilities. Other warning signs include unexplained “AI” labels, fixed demonstrations that cannot use buyer-supplied cases, material features shown only on a roadmap, no reliable export path, unclear ownership of configuration, and contracts that make source-data changes the customer’s problem without detection or reconciliation support.

Beware of a control that appears efficient because difficult cases, failed feeds, or aged work are excluded from reports. Ask how denominators are defined. Inspect raw counts alongside percentages.

Procurement implications and limitations

Score evidence, not presentation quality. A practical matrix gives separate weights to control fit, data and integration, governance, operations, security, implementation, commercial terms, and supplier viability. Record the evidence behind each score and any unresolved assumption. Require compliance and technology sign-off on mandatory requirements, security and privacy approval, model-risk review where relevant, procurement approval of terms, and an accountable executive decision.

This framework does not determine whether a product meets every legal obligation. Requirements vary by jurisdiction, institution, product, and risk assessment. A proof of concept uses a bounded sample, supplier evidence can become outdated, and future product changes may alter the result. The final record should state those limits, list conditions that must be satisfied before go-live, and assign owners for validation, migration, training, independent testing, and post-implementation review.