> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fuseai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Deduplicate contacts

> Find and merge duplicate contacts across your workspace: scan, review the groups, merge

Deduplication is asynchronous and runs in two explicit steps: a **scan** finds duplicate groups across your workspace's saved contacts and shows you exactly what each merge would produce, then a separate **merge** call applies the groups you choose. Nothing is changed until you start the merge, and neither step costs credits.

All three endpoints require a key with the `contacts` scope — see [Scopes](/concepts/scopes). The scan covers your team's saved contacts, Fuse and custom alike: everyone on your own lists and on your teammates' public lists, not one list.

## Start a scan

`POST /business/contacts/dedupe-scan` with the fields duplicates must match **exactly** on:

```bash theme={null}
curl -X POST "https://api.tryfuse.ai/api/v1/business/contacts/dedupe-scan" \
  -u "af_YOUR_KEY:" \
  -H "Content-Type: application/json" \
  -d '{ "matchFields": ["email"] }'
```

`matchFields` accepts any combination of `linkedinUrl`, `email`, `phone`, `fullName`, `jobTitle`, `industry`, `location`, `companyName`, `companyDomain` — with a floor: the selection must include **at least one unique identifier** (`linkedinUrl`, `email` or `phone`) **or at least three profile fields**. Anything weaker (say, `jobTitle` alone) would group unrelated people, and is rejected with `422 VALIDATION_FAILED`.

Values are compared after trimming and lowercasing. `email` and `phone` compare each contact's full set of addresses or numbers, so contacts that share only one of them are not grouped. A contact missing any selected field is left out of the scan.

The API responds `202` with a job id:

```json theme={null}
{
  "jobId": "68a1f5209c41d20014b3eb11"
}
```

A `202` means accepted, not finished. Hold on to `jobId` — it is both your handle for polling and the id the merge step consumes.

## Poll the scan

`GET /business/contacts/dedupe-jobs/{jobId}` returns the job's current state:

```bash theme={null}
curl "https://api.tryfuse.ai/api/v1/business/contacts/dedupe-jobs/68a1f5209c41d20014b3eb11" \
  -u "af_YOUR_KEY:"
```

`status` moves through the same lifecycle as every [long-running job](/concepts/long-running-work): `pending` → `processing` → `completed` or `failed`, and the terminal states are final. When the scan completes, `scanResult` carries **one page** of the duplicate groups — `pageNum` and `limit` page through them (default 50 per page, max 200, largest groups first), and `scanResult.pagination` says how many pages there are. Poll without paging params until `status` is `completed`, then page through the groups; a completed scan's full result can run to a megabyte, so re-fetching every group on every poll is wasted transfer.

```json theme={null}
{
  "dedupeJob": {
    "jobId": "68a1f5209c41d20014b3eb11",
    "kind": "scan",
    "status": "completed",
    "scanResult": {
      "matchFields": ["email"],
      "totalGroups": 1,
      "totalDuplicateContacts": 2,
      "totalContactsScanned": 4056,
      "truncated": false,
      "skippedOversizedGroups": 0,
      "pagination": { "pageNum": 1, "totalPages": 1, "totalRecords": 1 },
      "groups": [
        {
          "groupId": "3f6c2a9d1b8e4c7f",
          "matchedValues": { "email": "ada@brightpay.com" },
          "groupSize": 2,
          "contactIds": ["68a1f2c89c41d20014b3e955", "68a1f2c89c41d20014b3e956"],
          "lists": [{ "id": "68a1f20b9c41d20014b3e901", "name": "European fintech Series A" }],
          "listCount": 1,
          "mergePreview": {
            "winnerContactId": "68a1f2c89c41d20014b3e955",
            "firstName": "Ada",
            "lastName": "Nwosu",
            "jobTitle": "Head of Payroll",
            "linkedinUrl": "https://www.linkedin.com/in/ada-nwosu",
            "kind": "canonical",
            "companyName": "Brightpay",
            "companyDomain": "brightpay.com",
            "location": "Dublin, Leinster, Ireland",
            "primaryEmail": "ada@brightpay.com",
            "emailCount": 2,
            "emails": [
              { "email": "ada@brightpay.com", "status": "valid", "isPersonal": false },
              { "email": "ada.nwosu@gmail.com", "status": "valid", "isPersonal": true }
            ],
            "primaryPhone": "+353871234567",
            "phoneCount": 1,
            "phones": [{ "phoneNumber": "+353871234567", "status": "valid" }]
          },
          "contacts": [ ... ]
        }
      ]
    },
    "mergeResult": null,
    "error": null
  }
}
```

Each group is one set of contacts the scan believes are the same person:

* `contactIds` is every member; `contacts` previews the first 25 of them with their own fields, emails and phones.
* `mergePreview` is **what the surviving contact will look like** if you merge the group: the winner's identity (`winnerContactId`), empty fields filled from the other members, and the combined email/phone unions. `emails` and `phones` carry the first 10 entries each; `emailCount`/`phoneCount` are the full totals.
* `lists` names the lists the members live in, up to 10 of them, and `listCount` is the full total. Useful for judging blast radius before merging.
* `groupId` is the stable id the merge step selects by.

A scan stores at most 200 groups. `truncated: true` means more exist — merge what you have and scan again. Note the two totals side by side: `totalGroups` counts every group the scan *found* (it can exceed 200), while `pagination.totalRecords` counts the *stored*, pageable groups. Groups larger than 50 contacts are dropped and counted in `skippedOversizedGroups`; that many contacts matching exactly almost always means the field selection is too weak, not that one person has 50 records.

Scan results expire with the job after roughly 24 hours. Polling an unknown or expired `jobId` returns `404 JOB_NOT_FOUND`; start a new scan.

## Merge the groups

Review the groups, then `POST /business/contacts/dedupe-merge` with the scan's job id. Omit `groupIds` to merge every group, or pass a subset of `groupId` values to merge only those:

```bash theme={null}
curl -X POST "https://api.tryfuse.ai/api/v1/business/contacts/dedupe-merge" \
  -u "af_YOUR_KEY:" \
  -H "Content-Type: application/json" \
  -d '{ "scanJobId": "68a1f5209c41d20014b3eb11", "groupIds": ["3f6c2a9d1b8e4c7f"] }'
```

The API responds `202` with a new `jobId` — poll it at the same `GET /business/contacts/dedupe-jobs/{jobId}` endpoint. Within each group, the merge follows a fixed contract, the same one the in-app **Merge Duplicates** dialog applies:

* A Fuse contact beats a custom one as the survivor; among custom-only groups, the most recently created wins.
* The survivor's empty fields are filled from the other members.
* All emails and phones across the group are combined onto the survivor.
* Custom-column values are carried over.
* The losing contacts are archived, and every list membership they held is repointed to the survivor — no list loses a row.

When the merge job completes, `mergeResult` carries the totals:

```json theme={null}
{
  "dedupeJob": {
    "jobId": "68a1f6039c41d20014b3ec42",
    "kind": "merge",
    "status": "completed",
    "scanResult": null,
    "mergeResult": {
      "mergedGroups": 1,
      "archivedContacts": 1,
      "skippedGroups": 0
    },
    "error": null
  }
}
```

<Warning>
  Merging cannot be undone: the losing contacts are archived and their list memberships move to the survivor. Review `mergePreview` for every group you submit before starting the merge.
</Warning>

## Groups that will not merge

Two safety rules can make `mergedGroups` come out lower than the number of groups you submitted:

* **Conflicting Fuse identities.** Two Fuse contacts with different PDL ids or different LinkedIn URLs are never combined, even when the scan grouped them (they matched on your selected fields, but Fuse knows they are different people). Such groups are split into their consistent parts or skipped.
* **Stale groups.** A group whose members were already merged or archived by the time the job ran — by a teammate, or an earlier merge — is counted in `skippedGroups` rather than failing the job.

A `scanJobId` that is unknown or has expired answers `404 SCAN_NOT_FOUND`. Starting a merge against a scan that is still running answers `409 SCAN_NOT_COMPLETED`; poll the scan to `completed` first. A scan that found nothing answers `422 SCAN_HAS_NO_GROUPS` — there is nothing to merge, which is the good outcome.

## A compact end-to-end loop

```python theme={null}
import time
import requests

BASE = "https://api.tryfuse.ai/api/v1"
AUTH = ("af_YOUR_KEY", "")

def wait(job_id):
    delay = 2
    while True:
        job = requests.get(
            f"{BASE}/business/contacts/dedupe-jobs/{job_id}", auth=AUTH
        ).json()["dedupeJob"]
        if job["status"] in ("completed", "failed"):
            return job
        time.sleep(delay)
        delay = min(delay * 1.5, 15)

def fetch_all_groups(job_id):
    groups, page = [], 1
    while True:
        result = requests.get(
            f"{BASE}/business/contacts/dedupe-jobs/{job_id}",
            auth=AUTH, params={"pageNum": page, "limit": 50},
        ).json()["dedupeJob"]["scanResult"]
        groups.extend(result["groups"])
        if page >= result["pagination"]["totalPages"]:
            return groups
        page += 1

scan_id = requests.post(
    f"{BASE}/business/contacts/dedupe-scan", auth=AUTH,
    json={"matchFields": ["email"]},
).json()["jobId"]

scan = wait(scan_id)
groups = fetch_all_groups(scan_id) if scan["status"] == "completed" else []
if not groups:
    raise SystemExit("No duplicates found")

# Review before merging — e.g. only merge groups confined to one list.
group_ids = [g["groupId"] for g in groups if g["listCount"] <= 1]

merge_id = requests.post(
    f"{BASE}/business/contacts/dedupe-merge", auth=AUTH,
    json={"scanJobId": scan_id, "groupIds": group_ids},
).json()["jobId"]

merge = wait(merge_id)
print(merge["mergeResult"])
```

## Next steps

<CardGroup cols={2}>
  <Card title="Clean and enrich your CRM data" icon="arrows-rotate" href="/workflows/clean-and-enrich-crm-data">
    Import a CRM export, merge its duplicates, enrich it and sync it back in one workflow
  </Card>

  <Card title="API reference" icon="book" href="/api-reference">
    Full request and response schemas for the dedupe endpoints
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.