Data Validation in Clinical Data Management: Edit Checks, Queries and Cleaning

|Anders Mortin

Data validation is the part of clinical data management that checks the data itself is complete, consistent, and plausible — and resolves the problems it finds. It is the engine of data cleaning: the combination of automated checks, queries, and human review that turns raw entered data into data fit for analysis. Of all the conduct-stage activities, this is where data managers spend most of their time.

One clarification matters before anything else. In clinical trials, “validation” is used in two completely different senses. This guide is about data validation — checking that the data is correct. That is distinct from computer system validation, which proves that the EDC system itself works as intended; for that subject, see our guide to EDC validation requirements. Here, the focus is the data.

What is data validation in clinical data management?

Data validation is the process of confirming that collected data meets defined quality rules — that values are present where expected, fall within plausible ranges, agree with related data points, and are logically consistent. Where data fails a rule, the discrepancy is identified, queried, and resolved. Validation runs continuously through the conduct stage rather than as a single step, and its purpose is a clearly defined level of data quality: data fit for the decisions that depend on it.

Because not all data carries the same weight, cleaning is risk-based: effort concentrates on the critical data — safety and the primary endpoints — and once the data is fit for purpose, some residual issues in lower-impact fields are acceptable rather than chased to perfection.

It is worth repeating the distinction: data validation cleans the data, while computer system validation qualifies the system. A trial needs both, but they are different activities performed by different people for different reasons. The rest of this guide concerns data validation only.

Edit checks: the engine of data validation

Edit checks are the programmed and manual rules that test data against expectations. They are specified during set-up and run throughout conduct, flagging values that need attention. The common types are:

  • Missing data checks — flag required fields that have been left blank.
  • Range checks — flag values outside a plausible or permitted range, such as an impossible age or an out-of-range laboratory result.
  • Consistency or cross-field checks — flag values that contradict each other, such as a stop date earlier than a start date, or a finding inconsistent with a recorded condition.
  • Logical checks — flag combinations that should not occur given the protocol, such as a visit recorded out of sequence.

Most edit checks are automated and run in real time as data is entered, which is why their design during set-up matters so much. Not every automated check can run at the point of entry, though: some — spanning multiple visits, a patient's full set of adverse-event records, or data drawn from different systems — run as back-end (offline) checks in batch against submitted or consolidated data, with their findings then evaluated manually. The mechanics of how these checks are built and configured in an EDC system are covered in our complete guide to EDC systems; here the focus is what they are for and how they fit the cleaning process.

The data validation plan

The edit checks are not improvised — they are specified in a data validation plan (sometimes called a data validation specification). This document lists every check, what it tests, the message it raises, and how a flagged record should be handled. It is written and tested during set-up so that, when data starts to arrive, the cleaning rules are already defined, agreed, and validated. A clear data validation plan is what makes cleaning consistent and auditable rather than ad hoc.

Query management

When a check flags a value — or a reviewer spots a problem — a query is raised to the site to clarify or correct the data. Query management is the controlled workflow of raising, routing, answering, and closing these queries, with every step recorded in the audit trail. The aim is to resolve discrepancies at source, so the correction comes from the people who hold the original record rather than from an assumption. This reflects a core GCP principle: the investigator owns the trial data, and only the site can change it — a data manager or monitor cannot correct even an obvious error directly, but instead raises a query for the site to resolve. Tracking open and ageing queries is a core part of monitoring a trial's data-cleaning progress.

Discrepancy management and manual review

A discrepancy is any data point that does not meet expectations, whether flagged automatically or found by a person. Discrepancy management is the disciplined handling of these from identification to resolution. Not everything can be caught by a programmed rule, which is why automated checks are paired with manual data review — data managers and medical reviewers examining listings and patterns to find issues that no edit check anticipated. Source data review and source data verification complement both: in source data review (SDR) a monitor examines the site’s source records for completeness and consistency, while in source data verification (SDV) the source is compared directly against what was entered in the eCRF.  Together, automated checks, manual review, and source review and verification form the layered approach that produces a genuinely clean dataset.

Where data validation fits in the lifecycle

Data validation is a continuous conduct-stage activity. It begins as soon as data is entered and continues until cleaning is complete, working in parallel with reconciliation and medical coding. It feeds directly into the declaration of a clean file: a database cannot be declared clean while validation queries remain open. Completing data validation is therefore one of the prerequisites for database lock.

Frequently asked questions

What is data validation in clinical data management?

Data validation is the process of checking that collected trial data meets defined quality rules — complete, within range, consistent, and logical — and resolving any discrepancies found. It combines automated edit checks, query management, and manual review, and runs continuously through the conduct stage. Its goal is data of a defined quality, fit for analysis, and it must be complete before database lock.

What are edit checks in clinical data management?

Edit checks are programmed or manual rules that test data against expectations. Common types include missing-data checks, range checks for implausible values, consistency or cross-field checks for contradictory values, and logical checks for combinations that should not occur. Most are automated and run as data is entered, raising queries to the site when a value needs attention.

What is the difference between data validation and computer system validation?

Data validation checks that the data is correct — complete, consistent, and plausible — through edit checks and review. Computer system validation proves that the EDC system itself works as intended, through documented testing. They share the word “validation” but are different activities: one qualifies the data, the other qualifies the system. A trial needs both.

What is query management?

Query management is the controlled workflow of raising, routing, answering, and closing queries when data is flagged by a check or a reviewer. Every step is recorded in the audit trail, and the aim is to resolve discrepancies at source so corrections come from the holders of the original record. Tracking open and ageing queries is a key measure of cleaning progress.

What is a data validation plan?

A data validation plan, sometimes called a data validation specification, lists every edit check for a trial: what it tests, the message it raises, and how a flagged record should be handled. It is written and tested during set-up so that the cleaning rules are defined and agreed before data arrives, making the cleaning process consistent, complete, and auditable.

Learn data validation in context with TriTiCon

Data validation is the core of the conduct stage. TriTiCon's training covers it — edit checks, query management, and data cleaning — within the wider conduct process in The Clinical Data Management Conduct Process. You can explore the full TriTiCon course platform to see how data cleaning connects to the rest of the lifecycle.

Anders Mortin

Clinical Data Management Expert

TriTiCon delivers clinical data management training based on extensive hands-on experience from real clinical trials across sponsors, CROs, and life sciences organizations. The training is developed by industry professionals who work directly with clinical data, systems, documentation, and cross-functional trial teams.

30+
Years Experience
50+
Clinical Trials