How our DNA matching works
Every match on this site is found by comparing the raw DNA files you uploaded, position by position, and reporting the exact pieces of chromosome two people share. This page explains the method in one screen — and why the same two relatives can show different numbers at AncestryDNA, MyHeritage or GEDmatch.
1. What a match actually is
A raw DNA file is a long list of positions in your genome and the two alleles (A/C/G/T) you carry at each one. Matching two kits means:
- Walk both files together over the positions they both typed. Positions one file does not cover are ignored — they never break a segment.
- A shared position is one where the two people have at least one allele in common (you both carry an A, for example). A mismatch is an opposite call (AA vs GG).
- A segment is an unbroken run of shared positions. Its length is measured in centimorgans (cM), a unit of genetic distance from a recombination map — not in DNA base pairs.
- IBD2 (fully identical) is a run where both alleles match. Full siblings typically share IBD2 blocks; parent/child never do.
2. Our method, step by step
- The algorithm is the classic segment model that GEDmatch-style matching grew out of. It is deterministic: the same two files always give the same segments.
- A run needs ~200 shared positions to count (the algorithm's base threshold). Very short shared runs are ignored as coincidence.
- Isolated mismatches are bridged. A single opposite call with at least 150 shared positions to the next mismatch is treated as a read error and does not break the segment; a cluster of mismatches does.
- No-calls count as matches. If one file simply did not read a position, that is missing data, not disagreement — so it does not break the segment.
- Distance comes from the HapMap Phase II genetic map (GRCh37 / build 37). Its per-segment cM values line up closely with MyHeritage's, chromosome by chromosome.
- Segments from 6 cM up are reported; a pair is listed when its largest segment reaches 8 cM (the MyHeritage rule). We do not also demand a minimum number of SNPs — that would silently drop real cross-company segments, where the two chips overlap on fewer positions.
- SNP counts are raw shared counts — the number of positions both kits typed and matched. This is the same figure GEDmatch prints.
- X is compared separately (men carry one X). X matches are shown on the X filter, not added to the autosomal total.
3. How the big platforms differ
| AncestryDNA | MyHeritage | GEDmatch | RomanyDNA | |
|---|---|---|---|---|
| Chromosome browser | No — only total %/cM | Yes | Yes | Yes |
| Listed when | Not published (total cM based) | Largest segment ≥ 8 cM | Largest segment ≥ 7 cM and ≥ 700 SNPs | Largest segment ≥ 8 cM |
| Segments shown from | — | 6 cM | 7 cM (6 in one-to-one) | 6 cM |
| SNP counts shown | No | Internal figures (much higher than raw) | Raw shared counts | Raw shared counts (like GEDmatch) |
| Algorithm | Proprietary (+ Timber smoothing) | Proprietary (phasing/imputation) | Dynamic SNP threshold ~200, bunching, 500 kb gap breaks | Classic segment algorithm |
| Genetic map | Proprietary | Very close to HapMap Phase II | Their own map | HapMap Phase II |
| Database | Millions of testers | Millions of testers | Millions of uploads | Kits uploaded here only |
Thresholds above are the published/observable ones as of 2026. GEDmatch's one-to-one tool lets you change its settings; the one-to-many list uses the 7 cM / 700 SNP rule.
4. Why the same relatives show different numbers
If you compare the same two people at Ancestry, MyHeritage and GEDmatch you will get three different cM totals, segment counts and sometimes different segment boundaries. That is expected, not a mistake:
- Different genetic maps. Each company converts positions to cM with its own recombination map. The same segment can be 10% longer or shorter on one map than another.
- Different thresholds and gap rules. Companies differ in how many mismatches they tolerate, when they merge two nearby segments, and the minimum size they report.
- Phasing and imputation. Some platforms statistically fill in positions you were not tested for, or split DNA into maternal/paternal copies first. That changes where segments start and stop — and inflates their SNP counts.
- Different chips. An Ancestry kit and a MyHeritage kit type different positions, so only the overlap can be compared. Cross-company matches are therefore "thinner" in SNP counts than same-company ones at every platform.
Practical consequence: compare the largest segment, the number of segments and your tree against each other, not the absolute cM total from one site. A first cousin at ~850 cM on one site may read ~800–900 on another and still be the same relationship.
5. How to read our numbers
- Shared cM — the main closeness signal. Ranges: parent/child ≈ 3,400–3,600, full siblings ≈ 2,300–2,900, first cousins ≈ 550–1,300. These assume an outbred family.
- Largest segment — the strongest single piece of evidence; one long segment beats a pile of small ones.
- Segments — how many separate blocks. More, smaller blocks usually means a more distant relationship.
- IBD2 — fully identical blocks (double-shared) indicate close family or endogamy.
- Endogamy warning. In Roma and other endogamous families, cousins share more DNA than the standard ranges assume, so estimates read too close. Use the tree and triangulation to interpret them.
6. What we do not do
- We do not phase or impute your data for matching — only the DNA you actually have is compared.
- We do not report segments below 6 cM; below that level most shared runs are coincidence, not inheritance.
- We do not use a SNP-count floor to hide real cross-company matches (GEDmatch's 700-SNP rule is applied to its own dense data; on raw data it would delete genuine segments).
- We do not decide relationships from trees alone — cM totals, segment structure and your documented tree are all evidence.
Method attribution: segment detection uses the classic FF23utils segment model; genetic distances use the HapMap Phase II GRCh37 recombination map. Relationship ranges follow the widely used Shared cM Project reference tables.