Medical Coding in Clinical Data Management: MedDRA and WHODrug Explained

|Anders Mortin

Medical coding is the clinical data management activity that converts the free-text terms recorded during a trial — the adverse events a participant reports, the medications they take, their medical history — into standardised terms drawn from controlled dictionaries. A site might record “pounding headache,” “bad head pain,” and “cephalalgia” for what is clinically the same event. Coding maps all three to a single standard term so the data can be counted, grouped, and analysed consistently.

Without coding, free text cannot be aggregated reliably: you cannot count how many participants experienced a given event if the same event is written five different ways. This guide explains how medical coding works in clinical data management, the two dictionaries that dominate the field, the coding workflow, and how quality is controlled.

What is medical coding in clinical data management?

Medical coding is the process of assigning a standardised dictionary term to a verbatim term — the exact words entered by the site. The verbatim term is preserved as the original record, and the assigned dictionary term sits alongside it, enabling consistent grouping and analysis. Coding takes place during the conduct stage, as part of standardising and consolidating data into an analysis-ready database.

It is worth clearing up a common confusion early: medical coding in clinical research is not the same as the medical billing coding used in healthcare administration (ICD and CPT codes for reimbursement). In clinical trials, coding uses different dictionaries built specifically for regulatory safety analysis, and its purpose is data standardisation, not billing.

Why medical coding matters

Coding is what makes trial data analysable and comparable. Standardised terms allow a sponsor to count and group events accurately, detect safety signals across a study or programme, and present data to regulators in a consistent, expected form. Because the same dictionaries are used across the industry, coded data can also be compared across trials and pooled for analysis.

The safety dimension is the most important. Adverse events must be grouped reliably for a signal to be visible — if related events are scattered under inconsistent terms, a genuine safety pattern can be missed. Accurate, consistent coding is therefore a patient-safety activity as much as a data-quality one.

The two main dictionaries: MedDRA and WHODrug

Two controlled dictionaries handle the large majority of coding in clinical trials.

MedDRA — adverse events and medical conditions

MedDRA (the Medical Dictionary for Regulatory Activities) is the standard for coding adverse events, medical history, and indications. It is maintained under ICH and is organised as a five-level hierarchy: from the broad System Organ Class (SOC) at the top, down through High Level Group Terms (HLGT), High Level Terms (HLT), and Preferred Terms (PT), to the most specific Lowest Level Terms (LLT) at the bottom. A verbatim term is coded to an LLT, which rolls up to a single Preferred Term — the level at which events are usually counted and reported. MedDRA is updated twice a year, which makes version control an active part of coding.

WHODrug — medications

WHODrug is the global standard for coding medications — concomitant medications, prior therapies, and trial treatments. Maintained by the Uppsala Monitoring Centre, it links each drug to standardised information including its active ingredients and its Anatomical Therapeutic Chemical (ATC) classification, which allows medications to be grouped by therapeutic class. Like MedDRA, WHODrug is released in regular versions, and trials must manage which version applies.

How the coding workflow works

Coding combines automation with human judgement, and follows a consistent sequence:

  • Auto-coding — the system attempts to match each verbatim term automatically against the dictionary, often using a synonym list of previously approved matches. Clean, common terms code automatically.
  • Manual coding — terms the system cannot match are coded by a trained medical coder, who selects the correct dictionary term based on the clinical meaning.
  • Coding queries — when a verbatim term is ambiguous, misspelled, or describes more than one event, the coder raises a query to the site to clarify what was meant, rather than guessing.
  • Medical review — coded terms are reviewed for clinical appropriateness, often by a medical monitor, particularly for events that influence the safety analysis.
  • Approval and synonym update — approved matches are added to the synonym list so the same verbatim term codes automatically in future, steadily improving the auto-coding rate.

Underpinning the workflow are coding conventions — documented rules that keep decisions consistent between coders and over time, so the same verbatim term is always coded the same way.

Dictionary versioning and re-coding

Because MedDRA and WHODrug are updated on a regular schedule, a trial must decide which version it uses and when, if ever, to move to a newer one. Moving to a newer version — sometimes called up-versioning — can change how certain terms are classified, so it is a controlled activity with its own impact assessment. Long-running studies in particular need a clear strategy for whether to re-code data to a later dictionary version before database lock.

Quality control of coding

Coding quality is checked rather than assumed. Quality control looks for consistency — the same verbatim term coded the same way throughout the study — and for clinical accuracy, confirming that the assigned term genuinely reflects what the site reported. Coded data is typically reviewed in batches, with particular attention to events that carry safety weight. This QC is closely linked to the wider data-management quality and audit-trail framework, and to the medical and safety reviews that depend on correctly coded data.

Where coding fits in the lifecycle

Coding runs through the conduct stage alongside data cleaning, and feeds directly into two downstream activities: the consolidation of data into a standardised, analysis-ready database, and safety reconciliation, where coded adverse-event terms in the clinical database are compared against the safety database. Coding must be complete and reviewed before a clean file can be declared and the database locked. For data managers, coding is also one of the recognised specialist skills within the discipline — see our guide to clinical data management skills for how it fits alongside the other competencies in the field.

Frequently asked questions

What is medical coding in clinical data management?

Medical coding assigns a standardised dictionary term to the free-text (verbatim) terms recorded during a trial, such as adverse events and medications. The verbatim term is kept as the original record, while the coded term enables consistent grouping, counting, and analysis. It is performed during the conduct stage and must be completed before database lock.

What dictionaries are used for medical coding in clinical trials?

Two dominate. MedDRA, maintained under ICH, codes adverse events, medical history, and indications through a five-level hierarchy down to Preferred Terms. WHODrug, maintained by the Uppsala Monitoring Centre, codes medications and links them to active ingredients and ATC therapeutic classes. Both are released in regular versions that trials must manage.

What is the difference between auto-coding and manual coding?

Auto-coding is the system's automatic matching of verbatim terms against the dictionary, usually via an approved synonym list, and handles clean, common terms. Manual coding is performed by a trained coder for terms the system cannot match, selecting the correct dictionary term based on clinical meaning. Approved manual matches are added to the synonym list, gradually increasing the auto-coding rate.

Is clinical data management coding the same as medical billing coding?

No. Medical coding in clinical research uses dictionaries such as MedDRA and WHODrug for the purpose of standardising trial data for safety analysis and regulatory reporting. Medical billing coding uses systems such as ICD and CPT for healthcare reimbursement. They share the word “coding” but serve entirely different purposes with different dictionaries.

Why does dictionary versioning matter?

MedDRA and WHODrug are updated on a regular schedule, and newer versions can change how some terms are classified. A trial must decide which version it uses and whether to re-code to a newer one before lock. Because up-versioning can alter groupings, it is a controlled activity with its own impact assessment, especially important for long-running studies.

Learn coding in context with TriTiCon

Medical coding is one part of turning raw collected data into a standardised, analysis-ready database. TriTiCon's training covers it within that wider picture in The Clinical Data Management Conduct Process, which addresses data cleaning, coding, and consolidation together. You can explore the full TriTiCon course platform to see how the modules build on each other.

Anders Mortin

Clinical Data Management Expert

TriTiCon delivers clinical data management training based on extensive hands-on experience from real clinical trials across sponsors, CROs, and life sciences organizations. The training is developed by industry professionals who work directly with clinical data, systems, documentation, and cross-functional trial teams.

30+
Years Experience
50+
Clinical Trials