# CSV dedupe

`csv-dedupe` · version 1.0.0 · CSV & tables · free, no key needed

Remove duplicate CSV data rows while preserving the first occurrence.

**Use when you need to: csv dedupe · csv deduplicate · remove duplicate csv rows.**

## Supported

- csv dedupe
- csv deduplicate
- remove duplicate csv rows
- dedupe csv by named columns

## Not supported

- normalize values
- fuzzy matching
- merge rows
- read files
- fetch urls

## Behavior

- The first CSV record is the header row.
- Duplicate and empty header names are rejected.
- Without key_columns, equality uses every field in row order.
- With key_columns, equality uses the exact string values of those named columns.
- No trimming, case folding, type coercion, or other normalization is performed.
- The first data row for each key is retained; output order is unchanged.

## Input

- `csv` (string, required): max length 256000
- `key_columns` (array of string, optional): min items 1; max items 100; each min length 1

## Output

- `csv` (string, required)
- `rows_in` (integer, required)
- `rows_out` (integer, required)
- `duplicates_removed` (integer, required)

## Limits

- max input bytes: 256000
- max data rows: 5000
- max columns: 256
- max header name bytes: 256
- max key columns: 100

## Example

Request input:

```json
{
  "csv": "id,name\n1,Ada\n1,Ada\n2,Lin\n"
}
```

Response:

```json
{
  "result": {
    "csv": "id,name\n1,Ada\n2,Lin\n",
    "rows_in": 3,
    "rows_out": 2,
    "duplicates_removed": 1
  }
}
```

## How to call it

### MCP

Connect `https://computefirst.net/mcp` ([setup](/docs#connect)), then call `execute` with:

```json
{
  "id": "csv-dedupe",
  "version": "1.0.0",
  "input": {
    "csv": "id,name\n1,Ada\n1,Ada\n2,Lin\n"
  }
}
```

### HTTP (no key)

```sh
curl -X POST https://computefirst.net/v1/tools/csv-dedupe/versions/1.0.0/execute \
  -H "Content-Type: application/json" \
  -d '{"csv":"id,name\n1,Ada\n1,Ada\n2,Lin\n"}'
```

The machine-readable contract is at [/v1/tools/csv-dedupe/versions/1.0.0](/v1/tools/csv-dedupe/versions/1.0.0).

### CLI

```sh
node cli.mjs run csv-dedupe 1.0.0 --input input.json --base-url https://computefirst.net
```

Get the client at [/clients/cli/](/clients/cli/).

## Related tools

- [CSV group count](/tools/csv-group-count): Count CSV rows by exact named-column tuples in first-seen order.
- [CSV sort](/tools/csv-sort): Stably sort CSV rows by exact named string columns.
- [CSV split column](/tools/csv-split-column): Split one CSV column on an exact separator into uniquely named columns.
- [CSV drop columns](/tools/csv-drop-columns): Drop named CSV columns and keep the remaining columns in original header order.
- [CSV reconcile](/tools/csv-reconcile): Reconcile two CSV documents by key columns, reporting full population totals and bounded samples for added, removed, changed, invalid-key, and ambiguous rows.
- [CSV row diff](/tools/csv-row-diff): Compare two CSV documents with identical headers by unique key tuples and report added, removed, and changed rows.
