Quick start guide
Run your first harmonization end to end — from uploading your files to saving the harmonized dataset.
Before you begin
A session harmonizes one source dataset against one standard. To begin, you upload four files:
- Your unharmonized source dataset — the study-specific data to be transformed (
CSV) - The standard dataset that defines the target schema (
CSV) - The data dictionary for your source dataset (
PDF) - The data dictionary for the standard dataset (
PDF)
Before you begin
Make sure your datasets and data dictionaries are complete and correctly formatted before uploading. aha profiles variables exactly as they appear in the file — missing column descriptions may affect harmonization results.
Run your first harmonization
- Upload your four files. Add both datasets (CSV) and both data dictionaries (PDF), then click Next. Processing starts automatically and moves you into Data Profiling.
- Review profiling. Inspect each variable’s description, data type, and percentage of missing values — switching between the source and standard datasets via the tabs.
- Set the threshold and review the plan. Set the human-intervention threshold (recommended 60%), then review aha’s proposed decision, target, and transformations for each variable.
- Execution. aha applies the approved plan and shows the original and harmonized datasets side by side. No input is needed here.
- Validation. aha runs five quality checks on every harmonized variable. Address any that come back
FALSE. - Finalize & Save. When you’re satisfied, save the harmonized CSV to your workspace.
Reviewing as you go
Trust the scores
Every proposed plan carries a confidence score. Variables at or above your threshold are marked as not needing review — focus your attention on the ones flagged for human intervention.
Where to go next
Ready for the detail? Start with Data Profiling, the first module of the pipeline.
