aha logo aha™ Docs IGC Pharma v1.0 Download PDF

Quick start guide

Run your first harmonization end to end — from uploading your files to saving the harmonized dataset.

Before you begin

A session harmonizes one source dataset against one standard. To begin, you upload four files:

  • Your unharmonized source dataset — the study-specific data to be transformed (CSV)
  • The standard dataset that defines the target schema (CSV)
  • The data dictionary for your source dataset (PDF)
  • The data dictionary for the standard dataset (PDF)
Before you begin
Make sure your datasets and data dictionaries are complete and correctly formatted before uploading. aha profiles variables exactly as they appear in the file — missing column descriptions may affect harmonization results.

Run your first harmonization

  1. Upload your four files. Add both datasets (CSV) and both data dictionaries (PDF), then click Next. Processing starts automatically and moves you into Data Profiling.
  2. Review profiling. Inspect each variable’s description, data type, and percentage of missing values — switching between the source and standard datasets via the tabs.
  3. Set the threshold and review the plan. Set the human-intervention threshold (recommended 60%), then review aha’s proposed decision, target, and transformations for each variable.
  4. Execution. aha applies the approved plan and shows the original and harmonized datasets side by side. No input is needed here.
  5. Validation. aha runs five quality checks on every harmonized variable. Address any that come back FALSE.
  6. Finalize & Save. When you’re satisfied, save the harmonized CSV to your workspace.

Reviewing as you go

Trust the scores
Every proposed plan carries a confidence score. Variables at or above your threshold are marked as not needing review — focus your attention on the ones flagged for human intervention.

Where to go next

Ready for the detail? Start with Data Profiling, the first module of the pipeline.